awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
kyutai-labs avatar

kyutai-labs/delayed-streams-modeling

0
View on GitHub↗
2,955 stars·307 forks·Python·Apache-2.0·14 views

Delayed Streams Modeling

Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.

Features

  • Speech Processing - Real-time speech processing and modeling.
  • Speech Recognition - Real-time speech processing and modeling.

Star history

Star history chart for kyutai-labs/delayed-streams-modelingStar history chart for kyutai-labs/delayed-streams-modeling

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Delayed Streams Modeling

These projects share indexed features with Delayed Streams Modeling. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • openai/whisperopenai avatar

    openai/whisper

    102,828View on GitHub↗

    This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies

    Python
    View on GitHub↗102,828
  • qwenlm/qwen3-asrQwenLM avatar

    QwenLM/Qwen3-ASR

    1,603View on GitHub↗
    Python
    View on GitHub↗1,603
  • facebookresearch/omnilingual-asrfacebookresearch avatar

    facebookresearch/omnilingual-asr

    2,671View on GitHub↗

    Omnilingual-ASR is a multilingual automatic speech recognition framework and toolkit designed to transcribe audio across 1,600 languages. It provides a complete pipeline for converting speech to text, including a toolkit for fine-tuning pre-trained speech models to specific languages or datasets using custom training recipes. The system supports zero-shot speech recognition, allowing the model to predict text in unseen languages without extensive training data. It further enables few-shot language guidance through in-context examples and uses language codes to constrain transcription output t

    Python
    View on GitHub↗2,671
  • stepfun-ai/step-audio2S

    stepfun-ai/Step-Audio2

    0View on GitHub↗
    View on GitHub↗0
Compare all 30 related projects→

Frequently asked questions

What does kyutai-labs/delayed-streams-modeling do?

Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.

What are the main features of kyutai-labs/delayed-streams-modeling?

The main features of kyutai-labs/delayed-streams-modeling are: Speech Processing, Speech Recognition.

Which projects share features with kyutai-labs/delayed-streams-modeling?

Projects with overlapping indexed features include: qwenlm/qwen3-asr. xzf-thu/mega-asr. facebookresearch/omnilingual-asr — Omnilingual-ASR is a multilingual automatic speech recognition framework and toolkit designed to transcribe audio… openai/whisper — This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer… stepfun-ai/step-audio2. bytedance/megatts3 — MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English,…

Curated searches featuring Delayed Streams Modeling

Hand-picked collections where Delayed Streams Modeling appears.
  • Speech Synthesis and Recognition Models