awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

7 repositorios

Awesome GitHub RepositoriesTemporal Frame Alignment

Processes individual video frames to ensure precise temporal synchronization with corresponding audio segments.

Distinct from Frame Extractors: Focuses on the temporal alignment of frames to audio, rather than just sampling or extracting frames.

Explore 7 awesome GitHub repositories matching graphics & multimedia · Temporal Frame Alignment. Refine with filters or upvote what's useful.

Awesome Temporal Frame Alignment GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • rudrabha/wav2lipAvatar de Rudrabha

    Rudrabha/Wav2Lip

    13,045Ver en GitHub↗

    Wav2Lip is a deep learning lip sync model and neural talking head framework designed to synchronize the lip movements in a video to match a provided audio file. It functions as a computer vision lip synchronizer and speech-to-lip generator that maps speech patterns to visual mouth movements to produce realistic talking head videos. The system utilizes a framework for training and evaluating models that align audio and video frames. This includes the ability to train lip-sync models and visual discriminators using speech-to-lip datasets and evaluating the resulting synchronization accuracy thr

    Processes video sequences as individual frames to ensure perfect alignment with corresponding audio slices.

    Python
    Ver en GitHub↗13,045
  • jaywalnut310/vitsAvatar de jaywalnut310

    jaywalnut310/vits

    7,862Ver en GitHub↗

    This project is an end-to-end text-to-speech engine and deep learning voice synthesizer. It functions as a neural speech synthesis framework that converts written text directly into audio waveforms using a single neural network. The system implements an adversarial framework and a conditional variational autoencoder to generate high-fidelity artificial speech. It utilizes a generative adversarial network to ensure synthesized audio is indistinguishable from real human speech. The toolkit provides capabilities for neural speech synthesis, text-to-audio generation, and the training of custom v

    Automatically learns the alignment and duration between text characters and audio frames without external tools.

    Pythondeep-learningpytorchspeech-synthesis
    Ver en GitHub↗7,862
  • worldveil/dejavuAvatar de worldveil

    worldveil/dejavu

    6,764Ver en GitHub↗

    Dejavu is a Python audio fingerprinting library and recognition engine. It functions as a digital audio signature tool used to analyze sound waves and create unique identifiers for the purposes of audio search and retrieval. The project enables automatic music identification by matching live audio feeds or recorded clips against a database of fingerprints. It covers audio content matching and digital audio archiving to identify original source recordings from a stored collection. The system incorporates capabilities for generating audio fingerprints, identifying audio tracks, and recognizing

    Validates candidate matches by ensuring the temporal distance between fingerprints is consistent across the recording.

    Python
    Ver en GitHub↗6,764
  • ardour/ardourAvatar de Ardour

    Ardour/ardour

    5,057Ver en GitHub↗

    Ardour es una estación de trabajo de audio digital (DAW), mezclador de audio multipista y secuenciador MIDI. Funciona como un editor de audio no lineal y un host de plugins para ejecutar efectos e instrumentos de terceros. El sistema proporciona capacidades especializadas para la postproducción de audio mediante sincronización de fotogramas de vídeo, así como secuenciación de actuaciones en vivo para disparar clips y patrones en tiempo real. También soporta mezcla táctil mediante mapeo de superficies de control y configuración de controladores de hardware. El software cubre una amplia gama de necesidades de producción de audio, incluyendo grabación multipista, secuenciación y composición MIDI, mezcla profesional y exportación de audio multicanal. Su framework de procesamiento incluye soporte para plugins estándar de la industria y un sistema de enrutamiento de señales tipo matriz.

    Provides precise temporal alignment of audio segments with corresponding video frames for post-production scoring.

    C++audioc-plus-plusdaw
    Ver en GitHub↗5,057
  • plachtaa/vits-fast-fine-tuningAvatar de Plachtaa

    Plachtaa/VITS-fast-fine-tuning

    5,016Ver en GitHub↗

    VITS-fast-fine-tuning es un pipeline para adaptar modelos de síntesis de voz a voces objetivo específicas utilizando pequeños conjuntos de datos de audio. Funciona como una herramienta de adaptación rápida de hablante y un sintetizador de voz multilingüe capaz de generar audio hablado en diferentes idiomas. El sistema proporciona un framework para la conversión de voz muchos-a-muchos, transformando la identidad de un hablante en otro mientras se preserva el contenido lingüístico original. Permite la adaptación de una voz para texto-a-voz mediante el fine-tuning de un modelo pre-entrenado con clips de audio o fuentes de video. El proyecto cubre la síntesis de voz end-to-end y el procesamiento de audio, utilizando generación de formas de onda adversarias y búsqueda de alineación monótona para producir audio de alta fidelidad. Incorpora un predictor de duración estocástico para gestionar variaciones en el ritmo del habla y admite la transferencia de modelos pre-entrenados.

    Automatically learns the mapping between text characters and audio frames during the training process.

    Python
    Ver en GitHub↗5,016
  • gpac/gpacAvatar de gpac

    gpac/gpac

    3,205Ver en GitHub↗

    GPAC is an open-source multimedia framework built around a pluggable filter graph pipeline, where modular processing units called filters connect into a directed graph to handle media workflows. At its core, the framework centers all media packaging and manipulation on the ISO Base Media File Format (ISOBMFF), with specialized tools for reading, writing, fragmenting, and encrypting MP4 and related containers. It also provides a declarative scene graph composition system for describing interactive multimedia scenes using MPEG-4 BIFS, X3D, SVG, or VRML syntax, alongside a hardware-accelerated re

    Compares key-frame intervals and sync sample positions across files to detect misalignment before DASH packaging.

    Catsc3broadcastcenc
    Ver en GitHub↗3,205
  • intro-skipper/intro-skipperAvatar de intro-skipper

    intro-skipper/intro-skipper

    2,469Ver en GitHub↗

    Intro Skipper is a media server plugin and automated playback utility designed to identify and bypass television opening sequences. It functions as an automated content sequence skipper that detects repeated introduction segments in video files to improve viewing efficiency. The tool employs audio fingerprinting to analyze audio patterns during playback, comparing waveforms against known templates to trigger skip events. It allows for the management of playback preferences across multiple client devices to determine how these opening sequences are handled. The project covers automated media

    Analyzes time-stamped audio data to determine precise skip intervals for media files.

    C#jellyfinjellyfin-mediasegment-providerjellyfin-plugin
    Ver en GitHub↗2,469
  1. Home
  2. Graphics & Multimedia
  3. Image Processing & Editing
  4. Image Processing
  5. Frame Extractors
  6. Temporal Frame Alignment

Explorar subetiquetas

  • Audio Temporal Alignment1 sub-etiquetaVerifying that the time distance between fingerprints remains constant to validate a match. **Distinct from Temporal Frame Alignment:** Distinct from Temporal Frame Alignment: focuses on validating the relative timing of audio fingerprints rather than syncing frames to audio.
  • Key-Frame Alignment CheckersComparing key-frame intervals and sync sample positions across files to detect misalignment before packaging. **Distinct from Temporal Frame Alignment:** Distinct from Temporal Frame Alignment: focuses on checking alignment of key frames across multiple files for DASH packaging, not audio-video synchronization.