awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
vita-epfl avatar

vita-epfl/Stable-Video-Infinity

0
View on GitHub↗
2,493 stars·218 forks·Python·MIT·33 viewsstable-video-infinity.github.io/homepage↗

Stable Video Infinity

Stable-Video-Infinity is a video synthesis tool based on Stable Video Diffusion designed for creating long-form animations and consistent visual content. It serves as an AI video extension framework and a conditioned animation synthesizer capable of producing video sequences of arbitrary length.

The project enables infinite video extension by bypassing standard model duration constraints through an error-recycling loop. It supports conditioned animation synthesis using external inputs such as image streams, audio files, or skeletal motion data to guide the generation process.

The framework includes tools for diffusion model fine-tuning, allowing for custom video model training via specialized datasets and adapter-based techniques to create unique animation styles.

Features

  • Video Diffusion Models - Utilizes video diffusion models to synthesize content in a compressed latent space for efficiency.
  • Video Synthesis - Generates long-form video content conditioned on images, audio, and skeletal motion data.
  • Multi-Modal Conditioned Synthesis - Synthesizes long videos conditioned on diverse external inputs such as image streams, audio, and motion data.
  • Residual Recycling Loops - Bypasses standard duration constraints by feeding processing residuals back into the model.
  • Motion Conditioning - Integrates external audio and skeletal motion data to guide the animation synthesis process.
  • Arbitrary Duration Video Generators - Uses autoregressive prediction to generate video sequences of arbitrary length by predicting subsequent frames.
  • Infinite-Length Generators - Produces indefinitely long video sequences by bypassing standard model duration constraints.
  • Long-form Generation - Creates extended video sequences with smooth transitions and maintained visual consistency.
  • Adapter Fine-Tuning - Implements adapter-based layers to specialize animation styles without requiring full model retraining.
  • Temporal Attention - Employs temporal attention to maintain visual consistency and coherence across video frames.
  • Video Model Fine-Tuning - Provides tools for fine-tuning video models using specialized datasets and LoRA adapters.
  • Long Video Generation - Enables infinite-length video generation using error recycling techniques.

Star history

Star history chart for vita-epfl/stable-video-infinityStar history chart for vita-epfl/stable-video-infinity

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Stable Video Infinity

These projects share indexed features with Stable Video Infinity. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • ailab-cvc/videocrafterailab-cvc avatar

    ailab-cvc/videocrafter

    5,063View on GitHub↗

    Videocrafter is a latent diffusion model designed for AI video synthesis. It functions as both a text-to-video and image-to-video generation system, synthesizing high-quality video sequences from descriptive text prompts or static image inputs. The model utilizes a diffusion-based neural network to transform inputs into animated content, ensuring visual consistency and temporal coherence throughout the generated sequences. This allows for the creation of custom video clips and the animation of static images into fluid motion.

    Python
    View on GitHub↗5,063
  • lightricks/comfyui-ltxvideoLightricks avatar

    Lightricks/ComfyUI-LTXVideo

    3,840View on GitHub↗

    ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat

    Pythoncomfyuidiffusion-modelsdit
    View on GitHub↗3,840
  • hvision-nku/storydiffusionHVision-NKU avatar

    HVision-NKU/StoryDiffusion

    6,430View on GitHub↗

    StoryDiffusion is a generative AI system designed for consistent character image and video generation. It utilizes a pluggable cross-attention module to inject shared character representations into pretrained diffusion models, allowing for visual identity stability across multiple images and scenes without retraining the base model. The project features a video generation pipeline that produces temporally coherent sequences from text prompts or condition images. It employs a latent space motion interpolator to predict intermediate frames and semantic motion, enabling long-range video generati

    Jupyter Notebook
    View on GitHub↗6,430
  • antgroup/echomimicantgroup avatar

    antgroup/echomimic

    4,255View on GitHub↗

    EchoMimic is a multimodal human animation framework and diffusion-based video generator. It produces lifelike facial and semi-body animations of a reference image by synthesizing motion and appearance from various source data. The system enables portrait animation driven by audio, pose sequences, or driver videos. It features a landmark conditioning tool that allows for the precise control of facial movements by modifying specific landmark points. The framework covers multi-modal motion synthesis and the synchronization of reference images to match the physical movements of a target driver.

    Pythonaaai2025audio-driven-portrait-animationsaudio-driven-talking-face
    View on GitHub↗4,255
Compare all 30 related projects→

Frequently asked questions

What does vita-epfl/stable-video-infinity do?

Stable-Video-Infinity is a video synthesis tool based on Stable Video Diffusion designed for creating long-form animations and consistent visual content. It serves as an AI video extension framework and a conditioned animation synthesizer capable of producing video sequences of arbitrary length.

What are the main features of vita-epfl/stable-video-infinity?

The main features of vita-epfl/stable-video-infinity are: Video Diffusion Models, Video Synthesis, Multi-Modal Conditioned Synthesis, Residual Recycling Loops, Motion Conditioning, Arbitrary Duration Video Generators, Infinite-Length Generators, Long-form Generation.

Which projects share features with vita-epfl/stable-video-infinity?

Projects with overlapping indexed features include: ailab-cvc/videocrafter — Videocrafter is a latent diffusion model designed for AI video synthesis. It functions as both a text-to-video and… lightricks/comfyui-ltxvideo — ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It… hvision-nku/storydiffusion — StoryDiffusion is a generative AI system designed for consistent character image and video generation. It utilizes a… meituan-longcat/longcat-video — LongCat-Video is a collection of specialized models for video synthesis, featuring a large language model based… antgroup/echomimic — EchoMimic is a multimodal human animation framework and diffusion-based video generator. It produces lifelike facial… picsart-ai-research/text2video-zero — Text2Video-Zero is a text-to-video diffusion model and framework designed to synthesize temporally consistent video…