awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

4 个仓库

Awesome GitHub RepositoriesAudio Temporal Alignment

Verifying that the time distance between fingerprints remains constant to validate a match.

Distinct from Temporal Frame Alignment: Distinct from Temporal Frame Alignment: focuses on validating the relative timing of audio fingerprints rather than syncing frames to audio.

Explore 4 awesome GitHub repositories matching graphics & multimedia · Audio Temporal Alignment. Refine with filters or upvote what's useful.

Awesome Audio Temporal Alignment GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • jaywalnut310/vitsjaywalnut310 的头像

    jaywalnut310/vits

    7,862在 GitHub 上查看↗

    This project is an end-to-end text-to-speech engine and deep learning voice synthesizer. It functions as a neural speech synthesis framework that converts written text directly into audio waveforms using a single neural network. The system implements an adversarial framework and a conditional variational autoencoder to generate high-fidelity artificial speech. It utilizes a generative adversarial network to ensure synthesized audio is indistinguishable from real human speech. The toolkit provides capabilities for neural speech synthesis, text-to-audio generation, and the training of custom v

    Automatically learns the alignment and duration between text characters and audio frames without external tools.

    Pythondeep-learningpytorchspeech-synthesis
    在 GitHub 上查看↗7,862
  • worldveil/dejavuworldveil 的头像

    worldveil/dejavu

    6,764在 GitHub 上查看↗

    Dejavu is a Python audio fingerprinting library and recognition engine. It functions as a digital audio signature tool used to analyze sound waves and create unique identifiers for the purposes of audio search and retrieval. The project enables automatic music identification by matching live audio feeds or recorded clips against a database of fingerprints. It covers audio content matching and digital audio archiving to identify original source recordings from a stored collection. The system incorporates capabilities for generating audio fingerprints, identifying audio tracks, and recognizing

    Validates candidate matches by ensuring the temporal distance between fingerprints is consistent across the recording.

    Python
    在 GitHub 上查看↗6,764
  • plachtaa/vits-fast-fine-tuningPlachtaa 的头像

    Plachtaa/VITS-fast-fine-tuning

    5,016在 GitHub 上查看↗

    VITS-fast-fine-tuning 是一个使用小型音频数据集将语音合成模型适配到特定目标音色的流水线。它充当快速说话人适配工具和多语言语音合成器,能够生成跨不同语言的口语音频。 该系统提供了一个用于多对多语音转换的框架,在保留原始语言内容的同时转换说话人的身份。它允许通过使用音频片段或视频源微调预训练模型来适配文本转语音的音色。 该项目涵盖端到端语音合成和音频处理,利用对抗性波形生成和单调对齐搜索来产生高保真音频。它结合了随机持续时间预测器来管理说话节奏的变化,并支持预训练模型迁移。

    Automatically learns the mapping between text characters and audio frames during the training process.

    Python
    在 GitHub 上查看↗5,016
  • intro-skipper/intro-skipperintro-skipper 的头像

    intro-skipper/intro-skipper

    2,469在 GitHub 上查看↗

    Intro Skipper is a media server plugin and automated playback utility designed to identify and bypass television opening sequences. It functions as an automated content sequence skipper that detects repeated introduction segments in video files to improve viewing efficiency. The tool employs audio fingerprinting to analyze audio patterns during playback, comparing waveforms against known templates to trigger skip events. It allows for the management of playback preferences across multiple client devices to determine how these opening sequences are handled. The project covers automated media

    Analyzes time-stamped audio data to determine precise skip intervals for media files.

    C#jellyfinjellyfin-mediasegment-providerjellyfin-plugin
    在 GitHub 上查看↗2,469
  1. Home
  2. Graphics & Multimedia
  3. Image Processing & Editing
  4. Image Processing
  5. Frame Extractors
  6. Temporal Frame Alignment
  7. Audio Temporal Alignment

探索子标签

  • Monotonic Alignment SearchesAlgorithms that automatically learn the duration and alignment between text characters and audio frames. **Distinct from Audio Temporal Alignment:** Distinct from Audio Temporal Alignment: specifically learns character-to-frame duration mapping during training rather than validating existing fingerprints.