awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

115 repository-uri

Awesome GitHub RepositoriesVideo Generation

Models and frameworks for generating video content from text or image prompts.

Distinguishing note: No existing candidates for video generation; this is a distinct generative AI capability.

Explore 115 awesome GitHub repositories matching artificial intelligence & ml · Video Generation. Refine with filters or upvote what's useful.

Awesome Video Generation GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • comfyanonymous/comfyuiAvatar comfyanonymous

    comfyanonymous/ComfyUI

    117,322Vezi pe GitHub↗

    ComfyUI is a modular generative AI workflow orchestrator and node-based GUI for designing and executing complex diffusion model pipelines. It functions as both a visual interface for building generative logic graphs and a programmable backend API that exposes diffusion model operations for external integration. The system distinguishes itself through a graph-based execution model that supports differential workflow execution, re-running only modified nodes to reduce computation. It features dynamic model offloading to manage memory between system RAM and GPU VRAM and utilizes metadata-embedde

    Produces motion sequences and 3D assets by transforming text or images into high-quality video.

    Python
    Vezi pe GitHub↗117,322
  • heyputer/puterAvatar HeyPuter

    HeyPuter/puter

    42,318Vezi pe GitHub↗

    Puter is a browser-based desktop environment and cloud-native development platform that provides a virtualized graphical workspace. It enables developers to build and deploy full-stack web applications by integrating cloud storage, authentication, and serverless backend logic directly into the browser, eliminating the need for traditional server infrastructure. The platform distinguishes itself through a unified cloud storage layer and a distributed network runtime that facilitates peer-to-peer communication and cross-origin resource fetching. It features a sophisticated cross-window orchestr

    Creates videos using specific models by configuring frame rates, inference steps, and guidance scales.

    TypeScriptcloudcloud-oscloud-storage
    Vezi pe GitHub↗42,318
  • quantumnous/new-apiAvatar QuantumNous

    QuantumNous/new-api

    39,722Vezi pe GitHub↗

    This project is an AI model API gateway and proxy server designed to provide a unified interface for interacting with diverse artificial intelligence service providers. It functions as a centralized middleware platform that routes, load balances, and translates API requests across multiple models, enabling developers to access text, image, audio, and video generation capabilities through a single, standardized integration. The gateway distinguishes itself through comprehensive administrative and financial controls, including event-driven usage accounting, real-time token consumption tracking,

    Translates requests into video generation tasks for compatible models through a unified interface.

    Goai-gatewayclaudedeepseek
    Vezi pe GitHub↗39,722
  • google-research/google-researchAvatar google-research

    google-research/google-research

    38,139Vezi pe GitHub↗

    This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum computing, and large-scale scientific data analysis. It provides foundational frameworks for developing complex algorithmic systems, offering the necessary infrastructure for distributed training, computational graph execution, and high-performance model development. The project distinguishes itself by integrating specialized research domains with robust, privacy-preserving methodologies. It supports diverse scientific discovery through tools for quantum simulation, physics-informed

    Analyzes facial video clips using temporal shift neural networks to detect blood pulse fluctuations.

    Jupyter Notebookaimachine-learningresearch
    Vezi pe GitHub↗38,139
  • hpcaitech/open-soraAvatar hpcaitech

    hpcaitech/Open-Sora

    29,101Vezi pe GitHub↗

    Open-Sora is a video generation framework designed to produce cinematic sequences from text prompts and images. It functions as a generative system that transforms written descriptions or reference images into video content featuring realistic textures and lighting. The project includes a dedicated prompt engineering tool that uses large language models to expand simple user inputs into detailed descriptions. It also features a motion controller for adjusting movement intensity in generated sequences and evaluating motion levels in existing video files. The framework incorporates text-to-vid

    Provides a framework for generating high-quality cinematic video sequences from descriptive text prompts.

    Python
    Vezi pe GitHub↗29,101
  • sgl-project/sglangAvatar sgl-project

    sgl-project/sglang

    29,079Vezi pe GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Produces video content from text prompts and reference frames using dense or streaming pipelines.

    Pythonattentionblackwellcuda
    Vezi pe GitHub↗29,079
  • aidc-ai/pixelle-videoAvatar AIDC-AI

    AIDC-AI/Pixelle-Video

    23,403Vezi pe GitHub↗

    Pixelle-Video is a text-to-video automation platform and generation engine that converts text topics into complete videos with synchronized narration, images, and music. It functions as a modular system for producing short-form content, utilizing large language models to automate script composition, visual asset generation, and voiceover production. The platform features a node-based workflow orchestrator that allows the composition of custom generation pipelines by linking different AI models. It includes a dynamic video layout designer that uses HTML templates to define aspect ratios and vi

    Generates accompanying images or dynamic video clips using configurable workflows and style prompts.

    Pythonaigccomfyuiimage-generation
    Vezi pe GitHub↗23,403
  • vercel/aiAvatar vercel

    vercel/ai

    21,885Vezi pe GitHub↗

    This project is a comprehensive framework for building AI-powered applications, providing a unified toolkit for orchestrating language models, autonomous agents, and interactive user interfaces. It serves as a central library for managing the entire lifecycle of AI interactions, from initial prompt generation and model provider abstraction to complex, multi-step reasoning and tool execution. The framework distinguishes itself through its deep integration with frontend development, specifically by enabling generative user interfaces that render dynamic components directly from model outputs. I

    Interfaces with generative video models to synthesize content from text or image prompts.

    TypeScriptanthropicartificial-intelligencegemini
    Vezi pe GitHub↗21,885
  • anil-matcha/open-higgsfield-aiAvatar Anil-matcha

    Anil-matcha/Open-Higgsfield-AI

    20,529Vezi pe GitHub↗

    Open-Higgsfield-AI is a generative AI content studio and visual workflow orchestrator. It provides a unified interface for creating photorealistic images and videos, utilizing a node-based editor to chain multiple image, video, and audio models into automated content pipelines. The system functions as an AI video animation tool and local GPU inference engine, allowing users to run generative models on local hardware or remote servers. It includes specialized capabilities for audio-driven lip synchronization and cinematic camera controls to adjust virtual lens and focal settings. The platform

    Generates cinematic videos from text prompts or static images using generative AI and virtual camera controls.

    JavaScriptai-art-generatorai-image-generationai-video-generation
    Vezi pe GitHub↗20,529
  • titanwings/colleague-skillAvatar titanwings

    titanwings/colleague-skill

    19,817Vezi pe GitHub↗

    This project is a large language model persona simulation framework designed to distill individuals into AI personas using professional data and interpersonal context. It functions as a personal knowledge base ingestor and agent configuration manager, allowing for the creation of digital twins that reproduce a specific person's mental models, speaking styles, and professional workflows. The system utilizes a dual-model persona architecture that separates professional work skills from interpersonal personality traits. It distinguishes itself through a multimodal persona generator capable of pr

    Generates synthetic voice clones, images, and videos to provide visual and auditory dimensions to persona simulations.

    Python
    Vezi pe GitHub↗19,817
  • klingairesearch/liveportraitAvatar KlingAIResearch

    KlingAIResearch/LivePortrait

    17,830Vezi pe GitHub↗

    LivePortrait is a computer vision framework designed for portrait animation and generative video synthesis. It functions as a deep learning system that transfers facial expressions and head movements from a driving video source onto a static image or an existing portrait video, effectively decoupling the subject's identity from the dynamic motion patterns. The framework utilizes keypoint-based motion retargeting and implicit 3D latent representations to map movements across different subjects, including both human and animal portraits. By employing canonical motion normalization and feature-s

    Synthesizes realistic facial animations by mapping motion features from source media onto target subjects.

    Pythonface-animationimage-animationvideo-editing
    Vezi pe GitHub↗17,830
  • google-gemini/cookbookAvatar google-gemini

    google-gemini/cookbook

    17,418Vezi pe GitHub↗

    The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides designed for building applications with Google Gemini models. It serves as a central resource for developers to integrate multimodal generative artificial intelligence into their software, providing the necessary frameworks to manage model interactions, stateful workflows, and structured data extraction. The repository distinguishes itself by offering specialized toolkits for autonomous agent orchestration, enabling the construction of agents that can execute code, browse the web

    Creates high-fidelity videos from text prompts or reference images with support for custom aspect ratios and native audio.

    Jupyter Notebookgeminigemini-api
    Vezi pe GitHub↗17,418
  • lllyasviel/framepackAvatar lllyasviel

    lllyasviel/FramePack

    17,028Vezi pe GitHub↗

    FramePack is a neural video synthesis engine and generation framework designed to produce long, temporally consistent video sequences. It functions as a diffusion model optimizer, providing a suite of techniques to manage the computational demands of high-parameter video models while maintaining visual stability during extended generation tasks. The system distinguishes itself through a hierarchical approach to frame prediction, which plans distant anchor frames before filling in intermediate content to prevent cumulative temporal drift. By utilizing constant-length context compression and to

    Provides a comprehensive framework for training and deploying large-scale models capable of generating long, temporally consistent video sequences.

    Python
    Vezi pe GitHub↗17,028
  • rendercv/rendercvAvatar rendercv

    rendercv/rendercv

    16,955Vezi pe GitHub↗

    RenderCV is a command-line utility designed to transform structured YAML data into professionally typeset documents. By separating content from presentation, it allows users to maintain version-controlled resumes that are automatically rendered into high-quality PDF, HTML, and Markdown formats. The system leverages a specialized typesetting engine to ensure precise layout control and professional-grade typography. The project distinguishes itself through a schema-driven approach that enforces strict data validation, ensuring that input files are error-free before processing. Users can customi

    Displays real-time status and timing of the document generation process in the terminal.

    Pythoncvcv-buildercv-generator
    Vezi pe GitHub↗16,955
  • camenduru/stable-diffusion-webui-colabAvatar camenduru

    camenduru/stable-diffusion-webui-colab

    15,937Vezi pe GitHub↗

    This project provides a cloud-based notebook configuration for deploying a Stable Diffusion web interface. It functions as a specialized environment for image generation, incorporating a model trainer for fine-tuning weights and creating training datasets. The system emphasizes infrastructure persistence by saving software installations and model files to cloud storage, avoiding repetitive setups between sessions. It uses a tunnel-based interface to expose the web dashboard to a public URL for remote interaction. The project covers end-to-end AI workflows, including dataset preparation and t

    Synthesizes video sequences from textual prompts within a cloud environment.

    Jupyter Notebook
    Vezi pe GitHub↗15,937
  • vercel/vercelAvatar vercel

    vercel/vercel

    15,738Vezi pe GitHub↗

    Vercel is a cloud platform for building, deploying, and scaling web applications. It provides a unified infrastructure that automates the build process by detecting project frameworks and distributing static and dynamic content through a global content delivery network. The platform executes application logic using serverless functions that scale automatically based on real-time traffic demand. The platform distinguishes itself through a centralized AI gateway that proxies requests to multiple model providers, enabling standardized authentication, observability, and cost tracking. It supports

    Produces video output from text, image, or video inputs using advanced generative video models.

    TypeScriptclicloudcommand
    Vezi pe GitHub↗15,738
  • humanaigc/animateanyoneAvatar HumanAIGC

    HumanAIGC/AnimateAnyone

    14,774Vezi pe GitHub↗

    AnimateAnyone is an appearance-preserving video synthesizer designed for character animation from a single static image. It functions as a diffusion image-to-video generator that transforms a source image into a high-fidelity video sequence while maintaining consistent character identity, clothing, and visual details across all frames. The system enables video-driven character reenactment by transferring motions, facial expressions, and body movements from a reference video onto a static character. It employs pose-guided video generation to control movement via skeleton keypoints and pose sig

    Synthesizes high-fidelity video sequences of a character moving naturally using a single static image.

    Vezi pe GitHub↗14,774
  • bulletphysics/bullet3Avatar bulletphysics

    bulletphysics/bullet3

    14,243Vezi pe GitHub↗

    Bullet3 is a professional physics simulation engine designed for calculating rigid body, soft body, and collision dynamics within 3D environments and robotics applications. It functions as a computational framework for determining complex geometric intersections and contact manifolds between objects in simulated space. The library distinguishes itself through a distributed rendering framework that scales heavy graphical workloads and scene generation tasks across large clusters of machines. This capability enables the production of massive datasets by distributing complex scene generation acr

    Produces photo-realistic video scenes with automated annotations to support advanced machine learning research.

    C++computer-animationgame-developmentkinematics
    Vezi pe GitHub↗14,243
  • winfredy/sadtalkerAvatar Winfredy

    Winfredy/SadTalker

    13,919Vezi pe GitHub↗

    SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body

    Synthesizes expressive talking head videos by mapping audio signals to three-dimensional facial motion coefficients on static portrait images.

    Python
    Vezi pe GitHub↗13,919
  • opentalker/sadtalkerAvatar OpenTalker

    OpenTalker/SadTalker

    13,895Vezi pe GitHub↗

    SadTalker is an audio-driven talking head generator that produces synchronized speaking videos from a single source image and an input audio file. The system utilizes a deep learning framework to map speech signals to facial motion data, enabling the creation of lifelike digital avatars and animated characters. The project distinguishes itself by employing a three-dimensional morphable model to translate audio features into precise facial landmarks and head pose parameters. It integrates latent diffusion motion synthesis to generate naturalistic head movements and uses expression-aware textur

    Generates realistic videos of people speaking by mapping input audio and a single source image to precise facial motion data.

    Pythonaudio-driven-talking-facecvpr2023deep-fake
    Vezi pe GitHub↗13,895
Înapoi12345…6Înainte
  1. Home
  2. Artificial Intelligence & ML
  3. Video Generation

Explorează sub-etichetele

  • Arbitrary Duration Video GeneratorsGenerate videos of any length from sparse-frame input while maintaining temporal coherence. **Distinct from Video Generation:** Specific to generating arbitrarily long videos, distinct from general video generation which often produces fixed-length clips.
  • Generation Progress Trackers1 sub-tagTools for monitoring the lifecycle and completion status of asynchronous generative tasks. **Distinct from Video Generation:** Distinct from Video Generation: focuses on the status monitoring lifecycle rather than the generation process itself.
  • Generation StabilizationTechniques for reducing visual drifting in generated video sequences. **Distinct from Video Generation:** Distinct from Video Generation: focuses on the specific task of stabilizing long-form generation outputs.
  • Guided GenerationVideo generation combining text prompts with structural maps like pose, edge, or depth for layout control. **Distinct from Video Generation:** Specifically focuses on using structural maps for guidance, rather than general text or image prompts.
  • Identity-Preserving1 sub-tagSynthesizing video content that maintains the likeness of a specifically trained individual. **Distinct from Video Generation:** Distinct from Video Generation: specifically requires a trained identity model as a primary input for subject consistency.
  • Image-to-Video Generation2 sub-tag-uriCapabilities for synthesizing motion sequences using a reference image and text prompts as guidance. **Distinct from Video Generation:** Distinct from general Video Generation by specifically requiring a reference image as a primary input.
  • Infinite-Length Generators1 sub-tagVideo generation systems that create videos of any duration by extending a sequence frame-by-frame with independent noise-level denoising. **Distinct from Video Generation:** Distinct from Video Generation: generates videos of arbitrary length, not fixed-duration clips.
  • Lip-Synced1 sub-tagGenerating video content where facial movements are synchronized to an audio track. **Distinct from Video Generation:** Focuses on the specific task of lip-synced dubbing rather than general text-to-video generation.
  • Mixed-Asset SynthesisTechniques for generating videos by combining various media types like images and clips based on text descriptions. **Distinct from Video Generation:** Specifically focuses on the randomized pairing of different asset types rather than just generating a single video from a prompt.
  • Multi-Stage Refinement2 sub-tag-uriImproving visual fidelity and resolution using a two-stage structural-to-texture inference paradigm. **Distinct from Video Generation:** Specific to the multi-stage structural and texture refinement process of generative models
  • Multi-View Video SynthesisGenerating sequences of images from various camera angles to visualize reconstructed scenes. **Distinct from Video Generation:** Focuses on the multi-view visualization aspect of 3D reconstruction, rather than general generative video.
  • Personalized GenerationCapabilities for generating video content that maintains a specific identity or style across frames. **Distinct from Video Generation:** Focuses on identity preservation specifically within the context of video generation.
  • Progressive GenerationIterative frame prediction for extending video length incrementally. **Distinct from Video Generation:** Distinct from Video Generation: focuses on the progressive, iterative extension of existing sequences.
  • Prompt-Based Video Editors1 sub-tagTools for modifying existing video content using natural language instructions. **Distinct from Video Generation:** Focuses on generative editing of existing video rather than generating new video from scratch.
  • Reference-Based Video GeneratorsSystems that synthesize video content by combining character or scene references with text-based action prompts. **Distinct from Video Generation:** Distinct from general Video Generation: focuses specifically on using provided image references to define scene content.
  • ServicesAPIs for synthesizing video content from text or image inputs. **Distinct from Video Generation:** Distinct from Video Generation: focuses on the service-oriented API layer for video synthesis.
  • Streaming Generation1 sub-tagIncremental synthesis of video chunks for real-time file writing and immediate playback. **Distinct from Video Generation:** Focuses on the generative streaming process rather than the delivery of static video streams
  • Synthetic Video Generators1 sub-tagAutomated pipelines for producing photo-realistic video datasets with ground truth annotations. **Distinct from Video Generation:** Distinct from video generation: focuses on synthetic data production for machine learning rather than general video synthesis.
  • Temporal ExtensionsCapabilities for extending the duration of generated video clips to create longer motion sequences. **Distinct from Video Generation:** Distinct from general Video Generation: focuses specifically on increasing the length of existing clips rather than initial synthesis.
  • Temporal Sequence ExtensionCapabilities for extending the duration of existing video clips to produce longer, continuous sequences. **Distinct from Video Generation:** Distinct from general video generation by focusing on the temporal extension of existing content rather than creation from scratch
  • Text-to-Video Generators1 sub-tagSystems that synthesize video clips from text descriptions or reference images with configurable parameters. **Distinct from Video Generation:** Distinct from Video Generation: focuses on the text-to-video generation capability within Genkit, not general video generation frameworks.
  • VRAM-Constrained GeneratorsGenerates high-quality video content from text or images while managing memory constraints through model quantization and block offloading. **Distinct from Video Generation:** Distinct from Video Generation: specifically addresses VRAM constraints through quantization and offloading, not general video generation.
  • Video Clip Generators7 sub-tag-uriProduces video content from text prompts and reference frames using dense or streaming pipelines. **Distinct from Video Generation:** Focuses on clip-level generation rather than general video generation services.
  • World Simulation GeneratorsGenerates images, videos, synchronized sound, and action-conditioned rollouts from text, image, video, or action inputs for world simulation and synthetic data. **Distinct from Video Generation:** Distinct from Video Generation: focuses on generating world simulations with synchronized audio and action-conditioned rollouts for physical AI, not general video synthesis.