awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 个仓库

Awesome GitHub RepositoriesLatent Frame Transformations

Mapping audio features to latent visual representations for individual video frames.

Distinct from Video Frame Processing: Focuses on generative latent mapping for lip-sync rather than low-level GPU decoding or resizing.

Explore 2 awesome GitHub repositories matching graphics & multimedia · Latent Frame Transformations. Refine with filters or upvote what's useful.

Awesome Latent Frame Transformations GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • opentalker/video-retalkingOpenTalker 的头像

    OpenTalker/video-retalking

    7,256在 GitHub 上查看↗

    Video-retalking is an AI lip synchronization framework and talking head video editor designed to match the mouth movements of a subject in a video to a target audio track. It utilizes a deep learning pipeline to synchronize speech with video recordings. The system employs a two-stage generation process that separates coarse lip movement from high-resolution detail refinement. It incorporates identity-aware face refinement and expression template alignment to maintain photorealistic skin textures and ensure visual consistency across video frames. The toolset covers facial expression modificat

    Maps audio signals to latent representations that control the deformation of video frames for lip synchronization.

    Pythonlip-synchronizationsiggraph-asia-2022talking-head-videos
    在 GitHub 上查看↗7,256
  • tmelyralab/musetalkTMElyralab 的头像

    TMElyralab/MuseTalk

    5,327在 GitHub 上查看↗

    MuseTalk is a deep learning lip synchronization system designed to align video facial movements with audio tracks for high-fidelity video dubbing. It functions as an engine that matches facial expressions to audio input in real-time, enabling the modification of a speaker's lip movements to match new audio sources across different languages. The project features a distributed GPU training pipeline and a multi-stage processing workflow for refining the visual accuracy of synthetic speech. It distinguishes itself through the use of region-specific face masking and mouth openness control, which

    Translates audio features into frame-level visual transformations to ensure precise lip synchronization.

    Pythonlip-syncvirtualhumans
    在 GitHub 上查看↗5,327
  1. Home
  2. Graphics & Multimedia
  3. Video Frame Processing
  4. Latent Frame Transformations