awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

20 repository-uri

Awesome GitHub RepositoriesGenerative AI Tasks

Specific high-level tasks that involve the synthesis of new content from existing media inputs.

Explore 20 awesome GitHub repositories matching artificial intelligence & ml · Generative AI Tasks. Refine with filters or upvote what's useful.

Awesome Generative AI Tasks GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • comfy-org/comfyuiAvatar Comfy-Org

    Comfy-Org/ComfyUI

    117,227Vezi pe GitHub↗

    ComfyUI is a node-based generative AI orchestration engine designed for constructing, testing, and executing complex image and video synthesis pipelines. By utilizing a directed acyclic graph execution model, the platform allows users to build reproducible workflows through modular, interconnected processing blocks without requiring manual code implementation. It serves as both a local environment for high-performance model inference and a production-ready server for deploying generative capabilities. The platform distinguishes itself through its focus on workflow portability and extensibilit

    Transforms existing video footage into new visual outputs using reference-based generative models to maintain structural consistency.

    Pythonaicomfycomfyui
    Vezi pe GitHub↗117,227
  • kwaivgi/liveportraitAvatar KwaiVGI

    KwaiVGI/LivePortrait

    18,632Vezi pe GitHub↗

    LivePortrait is a deep learning framework for portrait animation that transfers facial expressions from a driving video to a static image. It functions as an AI motion retargeting tool, mapping movements between different identities while preserving the unique features of the source portrait. The system includes specialized capabilities for cross-species portrait animation, adapting human-centric models to non-human subjects and animals. It also features a motion template generator that converts driving videos into portable files to accelerate inference and protect the identity of the origina

    Implements generative retargeting of facial expressions and head movements between different identities.

    Python
    Vezi pe GitHub↗18,632
  • sczhou/codeformerAvatar sczhou

    sczhou/CodeFormer

    18,002Vezi pe GitHub↗

    CodeFormer is a deep learning framework designed for the restoration and enhancement of facial images and video sequences. It functions as a comprehensive processing engine capable of reconstructing high-quality facial features from degraded, blurry, or damaged inputs, while also providing tools for image upscaling and generative inpainting to fill missing or corrupted regions. The system distinguishes itself by utilizing a codebook-based quantization approach that maps input patches to high-quality facial representations, supported by transformer-based global modeling to ensure structural co

    Processes video sequences to improve facial quality across multiple frames using generative restoration models.

    Pythoncodebookcodeformerface-enhancement
    Vezi pe GitHub↗18,002
  • klingairesearch/liveportraitAvatar KlingAIResearch

    KlingAIResearch/LivePortrait

    17,830Vezi pe GitHub↗

    LivePortrait is a computer vision framework designed for portrait animation and generative video synthesis. It functions as a deep learning system that transfers facial expressions and head movements from a driving video source onto a static image or an existing portrait video, effectively decoupling the subject's identity from the dynamic motion patterns. The framework utilizes keypoint-based motion retargeting and implicit 3D latent representations to map movements across different subjects, including both human and animal portraits. By employing canonical motion normalization and feature-s

    Applies motion from a driving source to an existing portrait video to modify or retarget facial expressions and head movements.

    Pythonface-animationimage-animationvideo-editing
    Vezi pe GitHub↗17,830
  • lllyasviel/framepackAvatar lllyasviel

    lllyasviel/FramePack

    17,028Vezi pe GitHub↗

    FramePack is a neural video synthesis engine and generation framework designed to produce long, temporally consistent video sequences. It functions as a diffusion model optimizer, providing a suite of techniques to manage the computational demands of high-parameter video models while maintaining visual stability during extended generation tasks. The system distinguishes itself through a hierarchical approach to frame prediction, which plans distant anchor frames before filling in intermediate content to prevent cumulative temporal drift. By utilizing constant-length context compression and to

    Implements a neural engine that uses autoregressive processing and context compression to generate temporally consistent long-form video sequences.

    Python
    Vezi pe GitHub↗17,028
  • openai/baselinesAvatar openai

    openai/baselines

    16,733Vezi pe GitHub↗

    Baselines is a comprehensive suite of frameworks for reinforcement learning algorithm implementation, imitation learning, and training orchestration. It provides a library of standardized learning algorithms used to benchmark and replicate research results, alongside a deep learning policy framework for constructing neural network architectures such as multi-layer perceptrons, convolutional networks, and long short-term memory networks. The project includes a specialized imitation learning toolkit that enables agents to mimic expert behavior through behavior cloning and generative adversarial

    Captures and saves video clips of agents within simulation environments to monitor learning progress.

    Python
    Vezi pe GitHub↗16,733
  • aliaksandrsiarohin/first-order-modelAvatar AliaksandrSiarohin

    AliaksandrSiarohin/first-order-model

    15,003Vezi pe GitHub↗

    This project is a generative adversarial network designed for image animation and motion transfer. It functions as a computer vision framework that synthesizes video sequences by applying motion patterns extracted from a driving video onto a static source image. The model distinguishes itself by using a keypoint-based representation to decouple object appearance from temporal movement. By tracking structural deformations through learned latent coordinates, it performs motion retargeting and synthetic media production without requiring manual annotations or object-specific training data. The

    Transfers complex movement sequences from one video source to another subject to animate characters or faces.

    Jupyter Notebookdeep-learninggenerative-modelimage-animation
    Vezi pe GitHub↗15,003
  • rudrabha/wav2lipAvatar Rudrabha

    Rudrabha/Wav2Lip

    13,045Vezi pe GitHub↗

    Wav2Lip is a deep learning lip sync model and neural talking head framework designed to synchronize the lip movements in a video to match a provided audio file. It functions as a computer vision lip synchronizer and speech-to-lip generator that maps speech patterns to visual mouth movements to produce realistic talking head videos. The system utilizes a framework for training and evaluating models that align audio and video frames. This includes the ability to train lip-sync models and visual discriminators using speech-to-lip datasets and evaluating the resulting synchronization accuracy thr

    Matches a speaker's mouth movements to a new audio file using deep learning to maintain visual realism.

    Python
    Vezi pe GitHub↗13,045
  • wandb/wandbAvatar wandb

    wandb/wandb

    10,844Vezi pe GitHub↗

    Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow orchestration. It provides a comprehensive suite of tools for logging, visualizing, and versioning training metrics, model artifacts, and hyperparameter sweeps to ensure reproducibility across development cycles. The platform also functions as an observability tool for large language model applications, enabling the tracing of execution steps, token usage, and reasoning processes. The project distinguishes itself through its event-driven automation capabilities, which allow users

    Automatically saves and uploads video recordings of agent episodes from simulation environments.

    Pythonaicollaborationdata-science
    Vezi pe GitHub↗10,844
  • depthanything/depth-anything-v2Avatar DepthAnything

    DepthAnything/Depth-Anything-V2

    8,320Vezi pe GitHub↗

    Depth-Anything-V2 is a computer vision foundation model designed for general-purpose spatial understanding and depth perception. It functions as a monocular depth estimation model that predicts relative and absolute depth maps from single images or video sequences. The project provides specialized tools for both relative depth estimation and metric depth calculation, allowing for the determination of absolute physical distances in indoor and outdoor environments. It includes a video depth estimation framework that ensures temporal consistency across sequential frames to maintain stable depth

    Generates consistent depth maps across video frames to understand the three dimensional structure of moving scenes.

    Pythonmonocular-depth-estimation
    Vezi pe GitHub↗8,320
  • baowenbo/dainAvatar baowenbo

    baowenbo/DAIN

    8,311Vezi pe GitHub↗

    DAIN is a video frame synthesis engine and AI video upsampling tool designed to increase video playback smoothness. It functions as a computer vision model that synthesizes intermediate frames between existing images to transform low frame rate video into high frame rate content. The system utilizes depth-aware video frame interpolation to predict the motion of pixels between consecutive images. By analyzing spatial depth via depth maps, the tool generates new frames that account for occlusions and overlapping objects to create slow motion effects. The framework incorporates optical flow int

    Generates new video frames that account for spatial depth to avoid artifacts during movement.

    Python
    Vezi pe GitHub↗8,311
  • lipku/livetalkingAvatar lipku

    lipku/LiveTalking

    8,042Vezi pe GitHub↗

    LiveTalking is an interactive talking head engine and AI avatar management platform designed to synchronize synthetic speech with facial movements. It functions as a real-time orchestrator that connects large language models and text-to-speech services to neural-rendered digital humans. The project distinguishes itself through low-latency streaming capabilities and the ability to handle real-time conversational interruptions. It supports advanced audio-visual customization, including human voice cloning and the ability to drive avatar expressions using real-time webcam data. The platform cov

    Generate lip-synced digital human animations using neural rendering models to align visual speech with audio inputs.

    Pythonaigcdigihumandigital-human
    Vezi pe GitHub↗8,042
  • opentalker/video-retalkingAvatar OpenTalker

    OpenTalker/video-retalking

    7,256Vezi pe GitHub↗

    Video-retalking is an AI lip synchronization framework and talking head video editor designed to match the mouth movements of a subject in a video to a target audio track. It utilizes a deep learning pipeline to synchronize speech with video recordings. The system employs a two-stage generation process that separates coarse lip movement from high-resolution detail refinement. It incorporates identity-aware face refinement and expression template alignment to maintain photorealistic skin textures and ensure visual consistency across video frames. The toolset covers facial expression modificat

    Implements a deep learning pipeline to synchronize a subject's lip movements with a target audio track.

    Pythonlip-synchronizationsiggraph-asia-2022talking-head-videos
    Vezi pe GitHub↗7,256
  • tmelyralab/musetalkAvatar TMElyralab

    TMElyralab/MuseTalk

    5,327Vezi pe GitHub↗

    MuseTalk is a deep learning lip synchronization system designed to align video facial movements with audio tracks for high-fidelity video dubbing. It functions as an engine that matches facial expressions to audio input in real-time, enabling the modification of a speaker's lip movements to match new audio sources across different languages. The project features a distributed GPU training pipeline and a multi-stage processing workflow for refining the visual accuracy of synthetic speech. It distinguishes itself through the use of region-specific face masking and mouth openness control, which

    Modifies facial movements in video to match input audio across multiple languages while maintaining visual fidelity.

    Pythonlip-syncvirtualhumans
    Vezi pe GitHub↗5,327
  • badtobest/echomimicAvatar BadToBest

    BadToBest/EchoMimic

    4,258Vezi pe GitHub↗

    EchoMimic este un framework de animație a portretelor bazat pe audio și un generator video de difuzie latentă. Acesta transformă imaginile de referință statice în videoclipuri dinamice cu capete vorbitoare prin sincronizarea mișcărilor faciale cu piesele audio și driverele de mișcare. Sistemul funcționează ca un motor hibrid de sinteză a mișcării care combină input-urile audio și datele de postură. Utilizează un controler de mișcare a punctelor de reper faciale pentru a edita markerii de poziționare, permițând sincronizarea precisă și transferul de postură video-la-video. Conducta acoperă animația imagine-la-video prin difuzie latentă și condiționarea punctelor de reper faciale. Acest lucru permite animația portretelor condusă de audio, postură sau o combinație a ambelor surse de ghidare.

    Retargets facial expressions and head movements from a driver video to a reference portrait.

    Python
    Vezi pe GitHub↗4,258
  • ravendb/ravendbAvatar ravendb

    ravendb/ravendb

    3,961Vezi pe GitHub↗

    RavenDB is a multi-model NoSQL document database designed for high-performance, ACID-compliant data storage. It persists structured information as schema-flexible JSON documents and utilizes a unit-of-work session pattern to track entity changes and batch modifications into atomic transactions. The platform is built on a distributed architecture that supports horizontal scaling through sharding and ensures high availability via multi-node, master-to-master cluster replication. The database distinguishes itself through a self-optimizing query engine that automatically creates and maintains ind

    Embeds generative AI tasks directly into application workflows to automate data analysis and content generation.

    C#csharpdatabasedocument-database
    Vezi pe GitHub↗3,961
  • lightricks/comfyui-ltxvideoAvatar Lightricks

    Lightricks/ComfyUI-LTXVideo

    3,840Vezi pe GitHub↗

    ComfyUI-LTXVideo is a generative framework and ComfyUI custom node extension for synthesizing high-fidelity video. It utilizes a latent diffusion and transformer-based system to create cinematic clips from text, image, and audio inputs, providing a modular interface for precise control over subject behavior and temporal consistency. The tool distinguishes itself with production-grade capabilities, including the generation of High Dynamic Range video in linear formats such as ARRI LogC3. It supports multimodal synchronization for audio-driven animation and lip-syncing, and allows for the creat

    Synchronizes lip movements and scene pacing by integrating audio and visual data into a joint model.

    Pythoncomfyuidiffusion-modelsdit
    Vezi pe GitHub↗3,840
  • sandai-org/magi-1Avatar SandAI-org

    SandAI-org/MAGI-1

    3,711Vezi pe GitHub↗

    MAGI-1 is an autoregressive video generation model designed to synthesize high-resolution video sequences from text prompts and image references. It functions as a generative system for text-to-video, image-to-video, and video-to-video transformations. The model utilizes an autoregressive architecture that treats spatio-temporal patches as a sequence of discrete tokens to maintain temporal motion. It employs a variational autoencoder to compress the spatial and temporal dimensions of video data and uses distillation-based step scaling to allow for inference budget control. The system integra

    Uses an autoregressive synthesis engine to generate high-quality video frames with consistent temporal motion.

    Pythonautoregressivediffusion-modelsvideo-generation
    Vezi pe GitHub↗3,711
  • ali-vilab/vaceAvatar ali-vilab

    ali-vilab/VACE

    3,645Vezi pe GitHub↗

    VACE is a set of software tools and frameworks for reference-guided video generation, diffusion-based editing, and video-to-video translation. It provides utilities to produce new video content and modify existing sequences by using reference materials to guide visual style, subject matter, and composition. The framework enables video-to-video translation and synthesis, allowing for the update of visual styles and depth. It also functions as a video editor for modifying properties and content through reference-guided transformations. The system covers localized video editing and inpainting,

    Performs video-to-video synthesis by injecting original structural information into a diffusion process to maintain consistency.

    Pythonvideo-editingvideo-generation
    Vezi pe GitHub↗3,645
  • xlang-ai/osworldAvatar xlang-ai

    xlang-ai/OSWorld

    2,584Vezi pe GitHub↗

    OSWorld is an evaluation framework and multimodal agent benchmark designed to test the ability of large language models to complete complex tasks within virtualized operating system environments. It provides a virtualized desktop sandbox and a virtual machine orchestrator to deploy, snapshot, and reset cloud-based desktops, ensuring reproducible test states for AI agent interactions. The system distinguishes itself by providing an OS-level action space that translates model decisions into mouse clicks, keyboard inputs, and system commands. It employs a standardized interface to integrate vari

    Captures screenshots, action logs, and video recordings of agent episodes to verify task completion.

    Pythonagentartificial-intelligencebenchmark
    Vezi pe GitHub↗2,584
  1. Home
  2. Artificial Intelligence & ML
  3. Generative AI Resources
  4. Diffusion & Visual Synthesis Models
  5. Generative AI Tasks

Explorează sub-etichetele

  • AI Audio-to-Video SynchronizationDeep learning techniques for aligning visual mouth movements with new audio files while maintaining realism. **Distinct from Video-to-Video Synthesis:** Focuses on audio-driven lip synchronization specifically, rather than general video-to-video synthesis.
  • Depth-Aware Video SynthesisGenerating new video frames that use spatial depth maps to prevent visual artifacts during movement. **Distinct from Video-to-Video Synthesis:** Specifically targets depth-informed frame generation rather than general video-to-video generative transformations.
  • Video Depth AnalysisThe process of interpreting 3D structure from moving scenes via sequential depth maps. **Distinct from Depth-Aware Video Synthesis:** Focuses on analysis and understanding of existing video rather than synthesizing new frames.
  • Video-to-Video Synthesis3 sub-tag-uriTransforming input video sequences into new visual outputs using generative models.