awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

20 Repos

Awesome GitHub RepositoriesGenerative AI Tasks

Specific high-level tasks that involve the synthesis of new content from existing media inputs.

Explore 20 awesome GitHub repositories matching artificial intelligence & ml · Generative AI Tasks. Refine with filters or upvote what's useful.

Awesome Generative AI Tasks GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • comfy-org/comfyuiAvatar von Comfy-Org

    Comfy-Org/ComfyUI

    117,227Auf GitHub ansehen↗

    ComfyUI is a node-based generative AI orchestration engine designed for constructing, testing, and executing complex image and video synthesis pipelines. By utilizing a directed acyclic graph execution model, the platform allows users to build reproducible workflows through modular, interconnected processing blocks without requiring manual code implementation. It serves as both a local environment for high-performance model inference and a production-ready server for deploying generative capabilities. The platform distinguishes itself through its focus on workflow portability and extensibilit

    Transforms existing video footage into new visual outputs using reference-based generative models to maintain structural consistency.

    Pythonaicomfycomfyui
    Auf GitHub ansehen↗117,227
  • kwaivgi/liveportraitAvatar von KwaiVGI

    KwaiVGI/LivePortrait

    18,632Auf GitHub ansehen↗

    LivePortrait is a deep learning framework for portrait animation that transfers facial expressions from a driving video to a static image. It functions as an AI motion retargeting tool, mapping movements between different identities while preserving the unique features of the source portrait. The system includes specialized capabilities for cross-species portrait animation, adapting human-centric models to non-human subjects and animals. It also features a motion template generator that converts driving videos into portable files to accelerate inference and protect the identity of the origina

    Implements generative retargeting of facial expressions and head movements between different identities.

    Python
    Auf GitHub ansehen↗18,632
  • sczhou/codeformerAvatar von sczhou

    sczhou/CodeFormer

    18,002Auf GitHub ansehen↗

    CodeFormer is a deep learning framework designed for the restoration and enhancement of facial images and video sequences. It functions as a comprehensive processing engine capable of reconstructing high-quality facial features from degraded, blurry, or damaged inputs, while also providing tools for image upscaling and generative inpainting to fill missing or corrupted regions. The system distinguishes itself by utilizing a codebook-based quantization approach that maps input patches to high-quality facial representations, supported by transformer-based global modeling to ensure structural co

    Processes video sequences to improve facial quality across multiple frames using generative restoration models.

    Pythoncodebookcodeformerface-enhancement
    Auf GitHub ansehen↗18,002
  • klingairesearch/liveportraitAvatar von KlingAIResearch

    KlingAIResearch/LivePortrait

    17,830Auf GitHub ansehen↗

    LivePortrait is a computer vision framework designed for portrait animation and generative video synthesis. It functions as a deep learning system that transfers facial expressions and head movements from a driving video source onto a static image or an existing portrait video, effectively decoupling the subject's identity from the dynamic motion patterns. The framework utilizes keypoint-based motion retargeting and implicit 3D latent representations to map movements across different subjects, including both human and animal portraits. By employing canonical motion normalization and feature-s

    Applies motion from a driving source to an existing portrait video to modify or retarget facial expressions and head movements.

    Pythonface-animationimage-animationvideo-editing
    Auf GitHub ansehen↗17,830
  • lllyasviel/framepackAvatar von lllyasviel

    lllyasviel/FramePack

    17,028Auf GitHub ansehen↗

    FramePack is a neural video synthesis engine and generation framework designed to produce long, temporally consistent video sequences. It functions as a diffusion model optimizer, providing a suite of techniques to manage the computational demands of high-parameter video models while maintaining visual stability during extended generation tasks. The system distinguishes itself through a hierarchical approach to frame prediction, which plans distant anchor frames before filling in intermediate content to prevent cumulative temporal drift. By utilizing constant-length context compression and to

    Implements a neural engine that uses autoregressive processing and context compression to generate temporally consistent long-form video sequences.

    Python
    Auf GitHub ansehen↗17,028
  • openai/baselinesAvatar von openai

    openai/baselines

    16,733Auf GitHub ansehen↗

    Baselines is a comprehensive suite of frameworks for reinforcement learning algorithm implementation, imitation learning, and training orchestration. It provides a library of standardized learning algorithms used to benchmark and replicate research results, alongside a deep learning policy framework for constructing neural network architectures such as multi-layer perceptrons, convolutional networks, and long short-term memory networks. The project includes a specialized imitation learning toolkit that enables agents to mimic expert behavior through behavior cloning and generative adversarial

    Captures and saves video clips of agents within simulation environments to monitor learning progress.

    Python
    Auf GitHub ansehen↗16,733
  • aliaksandrsiarohin/first-order-modelAvatar von AliaksandrSiarohin

    AliaksandrSiarohin/first-order-model

    15,003Auf GitHub ansehen↗

    This project is a generative adversarial network designed for image animation and motion transfer. It functions as a computer vision framework that synthesizes video sequences by applying motion patterns extracted from a driving video onto a static source image. The model distinguishes itself by using a keypoint-based representation to decouple object appearance from temporal movement. By tracking structural deformations through learned latent coordinates, it performs motion retargeting and synthetic media production without requiring manual annotations or object-specific training data. The

    Transfers complex movement sequences from one video source to another subject to animate characters or faces.

    Jupyter Notebookdeep-learninggenerative-modelimage-animation
    Auf GitHub ansehen↗15,003
  • rudrabha/wav2lipAvatar von Rudrabha

    Rudrabha/Wav2Lip

    13,045Auf GitHub ansehen↗

    Wav2Lip is a deep learning lip sync model and neural talking head framework designed to synchronize the lip movements in a video to match a provided audio file. It functions as a computer vision lip synchronizer and speech-to-lip generator that maps speech patterns to visual mouth movements to produce realistic talking head videos. The system utilizes a framework for training and evaluating models that align audio and video frames. This includes the ability to train lip-sync models and visual discriminators using speech-to-lip datasets and evaluating the resulting synchronization accuracy thr

    Matches a speaker's mouth movements to a new audio file using deep learning to maintain visual realism.

    Python
    Auf GitHub ansehen↗13,045
  • wandb/wandbAvatar von wandb

    wandb/wandb

    10,844Auf GitHub ansehen↗

    Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow orchestration. It provides a comprehensive suite of tools for logging, visualizing, and versioning training metrics, model artifacts, and hyperparameter sweeps to ensure reproducibility across development cycles. The platform also functions as an observability tool for large language model applications, enabling the tracing of execution steps, token usage, and reasoning processes. The project distinguishes itself through its event-driven automation capabilities, which allow users

    Automatically saves and uploads video recordings of agent episodes from simulation environments.

    Pythonaicollaborationdata-science
    Auf GitHub ansehen↗10,844
  • depthanything/depth-anything-v2Avatar von DepthAnything

    DepthAnything/Depth-Anything-V2

    8,320Auf GitHub ansehen↗

    Depth-Anything-V2 is a computer vision foundation model designed for general-purpose spatial understanding and depth perception. It functions as a monocular depth estimation model that predicts relative and absolute depth maps from single images or video sequences. The project provides specialized tools for both relative depth estimation and metric depth calculation, allowing for the determination of absolute physical distances in indoor and outdoor environments. It includes a video depth estimation framework that ensures temporal consistency across sequential frames to maintain stable depth

    Generates consistent depth maps across video frames to understand the three dimensional structure of moving scenes.

    Pythonmonocular-depth-estimation
    Auf GitHub ansehen↗8,320
  • baowenbo/dainAvatar von baowenbo

    baowenbo/DAIN

    8,311Auf GitHub ansehen↗

    DAIN is a video frame synthesis engine and AI video upsampling tool designed to increase video playback smoothness. It functions as a computer vision model that synthesizes intermediate frames between existing images to transform low frame rate video into high frame rate content. The system utilizes depth-aware video frame interpolation to predict the motion of pixels between consecutive images. By analyzing spatial depth via depth maps, the tool generates new frames that account for occlusions and overlapping objects to create slow motion effects. The framework incorporates optical flow int

    Generates new video frames that account for spatial depth to avoid artifacts during movement.

    Python
    Auf GitHub ansehen↗8,311
  • lipku/livetalkingAvatar von lipku

    lipku/LiveTalking

    8,042Auf GitHub ansehen↗

    LiveTalking is an interactive talking head engine and AI avatar management platform designed to synchronize synthetic speech with facial movements. It functions as a real-time orchestrator that connects large language models and text-to-speech services to neural-rendered digital humans. The project distinguishes itself through low-latency streaming capabilities and the ability to handle real-time conversational interruptions. It supports advanced audio-visual customization, including human voice cloning and the ability to drive avatar expressions using real-time webcam data. The platform cov

    Generate lip-synced digital human animations using neural rendering models to align visual speech with audio inputs.

    Pythonaigcdigihumandigital-human
    Auf GitHub ansehen↗8,042
  • opentalker/video-retalkingAvatar von OpenTalker

    OpenTalker/video-retalking

    7,256Auf GitHub ansehen↗

    Video-retalking is an AI lip synchronization framework and talking head video editor designed to match the mouth movements of a subject in a video to a target audio track. It utilizes a deep learning pipeline to synchronize speech with video recordings. The system employs a two-stage generation process that separates coarse lip movement from high-resolution detail refinement. It incorporates identity-aware face refinement and expression template alignment to maintain photorealistic skin textures and ensure visual consistency across video frames. The toolset covers facial expression modificat

    Implements a deep learning pipeline to synchronize a subject's lip movements with a target audio track.

    Pythonlip-synchronizationsiggraph-asia-2022talking-head-videos
    Auf GitHub ansehen↗7,256
  • tmelyralab/musetalkAvatar von TMElyralab

    TMElyralab/MuseTalk

    5,327Auf GitHub ansehen↗

    MuseTalk is a deep learning lip synchronization system designed to align video facial movements with audio tracks for high-fidelity video dubbing. It functions as an engine that matches facial expressions to audio input in real-time, enabling the modification of a speaker's lip movements to match new audio sources across different languages. The project features a distributed GPU training pipeline and a multi-stage processing workflow for refining the visual accuracy of synthetic speech. It distinguishes itself through the use of region-specific face masking and mouth openness control, which

    Modifies facial movements in video to match input audio across multiple languages while maintaining visual fidelity.

    Pythonlip-syncvirtualhumans
    Auf GitHub ansehen↗5,327
  • badtobest/echomimicAvatar von BadToBest

    BadToBest/EchoMimic

    4,258Auf GitHub ansehen↗

    EchoMimic is an audio-driven portrait animation framework and latent diffusion video generator. It transforms static reference images into dynamic talking head videos by synchronizing facial movements with audio tracks and motion drivers. The system functions as a hybrid motion synthesis engine that combines audio inputs and pose data. It utilizes a facial landmark motion controller to edit positioning markers, enabling precise synchronization and video-to-video pose transfer. The pipeline covers image-to-video animation through latent diffusion and facial landmark conditioning. This allows

    Retargets facial expressions and head movements from a driver video to a reference portrait.

    Python
    Auf GitHub ansehen↗4,258
  • ravendb/ravendbAvatar von ravendb

    ravendb/ravendb

    3,961Auf GitHub ansehen↗

    RavenDB is a multi-model NoSQL document database designed for high-performance, ACID-compliant data storage. It persists structured information as schema-flexible JSON documents and utilizes a unit-of-work session pattern to track entity changes and batch modifications into atomic transactions. The platform is built on a distributed architecture that supports horizontal scaling through sharding and ensures high availability via multi-node, master-to-master cluster replication. The database distinguishes itself through a self-optimizing query engine that automatically creates and maintains ind

    Embeds generative AI tasks directly into application workflows to automate data analysis and content generation.

    C#csharpdatabasedocument-database
    Auf GitHub ansehen↗3,961
  • lightricks/comfyui-ltxvideoAvatar von Lightricks

    Lightricks/ComfyUI-LTXVideo

    3,840Auf GitHub ansehen↗

    ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat

    Synchronizes lip movements and scene pacing by integrating audio and visual data into a joint model.

    Pythoncomfyuidiffusion-modelsdit
    Auf GitHub ansehen↗3,840
  • sandai-org/magi-1Avatar von SandAI-org

    SandAI-org/MAGI-1

    3,711Auf GitHub ansehen↗

    MAGI-1 is an autoregressive video generation model designed to synthesize high-resolution video sequences from text prompts and image references. It functions as a generative system for text-to-video, image-to-video, and video-to-video transformations. The model utilizes an autoregressive architecture that treats spatio-temporal patches as a sequence of discrete tokens to maintain temporal motion. It employs a variational autoencoder to compress the spatial and temporal dimensions of video data and uses distillation-based step scaling to allow for inference budget control. The system integra

    Uses an autoregressive synthesis engine to generate high-quality video frames with consistent temporal motion.

    Pythonautoregressivediffusion-modelsvideo-generation
    Auf GitHub ansehen↗3,711
  • ali-vilab/vaceAvatar von ali-vilab

    ali-vilab/VACE

    3,645Auf GitHub ansehen↗

    VACE is a set of software tools and frameworks for reference-guided video generation, diffusion-based editing, and video-to-video translation. It provides utilities to produce new video content and modify existing sequences by using reference materials to guide visual style, subject matter, and composition. The framework enables video-to-video translation and synthesis, allowing for the update of visual styles and depth. It also functions as a video editor for modifying properties and content through reference-guided transformations. The system covers localized video editing and inpainting,

    Performs video-to-video synthesis by injecting original structural information into a diffusion process to maintain consistency.

    Pythonvideo-editingvideo-generation
    Auf GitHub ansehen↗3,645
  • xlang-ai/osworldAvatar von xlang-ai

    xlang-ai/OSWorld

    2,584Auf GitHub ansehen↗

    OSWorld is an evaluation framework and multimodal agent benchmark designed to test the ability of large language models to complete complex tasks within virtualized operating system environments. It provides a virtualized desktop sandbox and a virtual machine orchestrator to deploy, snapshot, and reset cloud-based desktops, ensuring reproducible test states for AI agent interactions. The system distinguishes itself by providing an OS-level action space that translates model decisions into mouse clicks, keyboard inputs, and system commands. It employs a standardized interface to integrate vari

    Captures screenshots, action logs, and video recordings of agent episodes to verify task completion.

    Pythonagentartificial-intelligencebenchmark
    Auf GitHub ansehen↗2,584
  1. Home
  2. Artificial Intelligence & ML
  3. Generative AI Resources
  4. Diffusion & Visual Synthesis Models
  5. Generative AI Tasks

Unter-Tags erkunden

  • AI Audio-to-Video SynchronizationDeep learning techniques for aligning visual mouth movements with new audio files while maintaining realism. **Distinct from Video-to-Video Synthesis:** Focuses on audio-driven lip synchronization specifically, rather than general video-to-video synthesis.
  • Depth-Aware Video SynthesisGenerating new video frames that use spatial depth maps to prevent visual artifacts during movement. **Distinct from Video-to-Video Synthesis:** Specifically targets depth-informed frame generation rather than general video-to-video generative transformations.
  • Video Depth AnalysisThe process of interpreting 3D structure from moving scenes via sequential depth maps. **Distinct from Depth-Aware Video Synthesis:** Focuses on analysis and understanding of existing video rather than synthesizing new frames.
  • Video-to-Video Synthesis3 Sub-TagsTransforming input video sequences into new visual outputs using generative models.