awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 dépôts

Awesome GitHub RepositoriesLatent Frame Transformations

Mapping audio features to latent visual representations for individual video frames.

Distinct from Video Frame Processing: Focuses on generative latent mapping for lip-sync rather than low-level GPU decoding or resizing.

Explore 2 awesome GitHub repositories matching graphics & multimedia · Latent Frame Transformations. Refine with filters or upvote what's useful.

Awesome Latent Frame Transformations GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • opentalker/video-retalkingAvatar de OpenTalker

    OpenTalker/video-retalking

    7,256Voir sur GitHub↗

    Video-retalking is an AI lip synchronization framework and talking head video editor designed to match the mouth movements of a subject in a video to a target audio track. It utilizes a deep learning pipeline to synchronize speech with video recordings. The system employs a two-stage generation process that separates coarse lip movement from high-resolution detail refinement. It incorporates identity-aware face refinement and expression template alignment to maintain photorealistic skin textures and ensure visual consistency across video frames. The toolset covers facial expression modificat

    Maps audio signals to latent representations that control the deformation of video frames for lip synchronization.

    Pythonlip-synchronizationsiggraph-asia-2022talking-head-videos
    Voir sur GitHub↗7,256
  • tmelyralab/musetalkAvatar de TMElyralab

    TMElyralab/MuseTalk

    5,327Voir sur GitHub↗

    MuseTalk is a deep learning lip synchronization system designed to align video facial movements with audio tracks for high-fidelity video dubbing. It functions as an engine that matches facial expressions to audio input in real-time, enabling the modification of a speaker's lip movements to match new audio sources across different languages. The project features a distributed GPU training pipeline and a multi-stage processing workflow for refining the visual accuracy of synthetic speech. It distinguishes itself through the use of region-specific face masking and mouth openness control, which

    Translates audio features into frame-level visual transformations to ensure precise lip synchronization.

    Pythonlip-syncvirtualhumans
    Voir sur GitHub↗5,327
  1. Home
  2. Graphics & Multimedia
  3. Video Frame Processing
  4. Latent Frame Transformations