awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

9 مستودعات

Awesome GitHub RepositoriesInference Acceleration

Optimization techniques to reduce the computational cost and time of diffusion model sampling.

Distinct from Diffusion Models: Focuses on sampling speed and step reduction specifically for diffusion processes.

Explore 9 awesome GitHub repositories matching artificial intelligence & ml · Inference Acceleration. Refine with filters or upvote what's useful.

Awesome Inference Acceleration GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • instantx-research/instantidالصورة الرمزية لـ instantX-research

    instantX-research/InstantID

    11,955عرض على GitHub↗

    InstantID is a diffusion-based identity preservation framework designed for zero-shot image generation. It allows for the synthesis of images featuring a specific person's facial identity using a single reference photo without requiring additional model training or fine-tuning. The project distinguishes itself through the use of consistency model distillation to accelerate inference, reducing the number of steps needed to produce high-quality results. It combines identity-preserving feature extraction with multi-modal prompt integration to merge visual embeddings from a reference image with t

    Reduces the time and computational steps required for high-quality image generation using consistency models.

    Python
    عرض على GitHub↗11,955
  • nvlabs/sanaالصورة الرمزية لـ NVlabs

    NVlabs/Sana

    8,310عرض على GitHub↗

    Sana is a framework for high-resolution image and video synthesis based on a linear diffusion transformer. It provides a toolkit for the training, fine-tuning, and execution of text-to-image and text-to-video models, as well as a video generative world model capable of simulating physical environments with precise spatial control. The project is distinguished by its use of linear complexity layers to handle high resolutions and its support for long-form, minute-length video generation in real time. It implements a two-stage inference paradigm that separates structural generation from visual t

    Accelerates model exploration by sampling large numbers of candidates using low-precision quantization to filter for high-contrast seeds.

    Python
    عرض على GitHub↗8,310
  • open-mmlab/mmagicالصورة الرمزية لـ open-mmlab

    open-mmlab/mmagic

    7,434عرض على GitHub↗

    mmagic is a multimodal training pipeline and framework for generative AI, focusing on visual synthesis and restoration. It provides the infrastructure to build and train models for tasks such as text-to-image and text-to-video generation, 3D-aware content synthesis, and high-fidelity image translation using diffusion models and generative adversarial networks. The project distinguishes itself through specialized capabilities for generative model personalization, including techniques for fine-tuning subjects and styles. It also supports advanced visual manipulations such as latent space interp

    Accelerates diffusion model sampling by merging redundant tokens in the vision transformer.

    Jupyter Notebookaigccomputer-visiondeep-learning
    عرض على GitHub↗7,434
  • moonintheriver/diffsingerالصورة الرمزية لـ MoonInTheRiver

    MoonInTheRiver/DiffSinger

    4,804عرض على GitHub↗

    DiffSinger هو مركب صوتي للذكاء الاصطناعي ومولد صوت عصبي مصمم لإنتاج غناء وكلام عالي الدقة. يعمل كنظام تحويل النص إلى كلام وأداة تركيب صوت غنائي قائمة على الانتشار (Diffusion) تحول النص وطبقة الصوت إلى صوت مسموع. يستخدم النظام آلية انتشار ضحلة وتحسين ضوضاء تكراري لإنتاج عروض صوتية واقعية. ويدمج إضافات أخذ عينات متخصصة ومحلات عددية لتسريع الاستدلال وتقليل الوقت المطلوب لتوليد أصوات اصطناعية. يغطي المشروع النمذجة الصوتية، وتركيب مخطط ميل الطيفي (Mel-spectrogram)، وإعادة بناء الموكل العصبي (Neural vocoder) لتحويل النص إلى أشكال موجية صوتية في النطاق الزمني. كما يتضمن قدرات لتحسين الصوت الاصطناعي لتحسين الجودة الصوتية للتسجيلات.

    Optimizes inference speed by employing specialized numerical solvers to reduce the number of diffusion iterations.

    Pythonaaai2022diffusion-modeldiffusion-speedup
    عرض على GitHub↗4,804
  • tencent-hunyuan/hunyuanvideo-1.5الصورة الرمزية لـ Tencent-Hunyuan

    Tencent-Hunyuan/HunyuanVideo-1.5

    4,440عرض على GitHub↗

    HunyuanVideo-1.5 is a video generation foundation model and text-to-video diffusion framework. It utilizes a latent video diffusion model and a spatio-temporal transformer architecture to generate high-definition video sequences from text descriptions and images. The project enables cinematic camera control for directing pans and tilts and provides image-to-video animation capabilities. It supports visual style adaptation through low-rank adaptation tuning and uses a language model for prompt refinement to improve visual alignment. The model covers high-resolution video upscaling via a super

    Reduces video generation time through step distillation, cache inference, and sparse attention techniques.

    Pythonimage-to-videotext-to-videovideo-generation
    عرض على GitHub↗4,440
  • tencent-hunyuan/hunyuanditالصورة الرمزية لـ Tencent-Hunyuan

    Tencent-Hunyuan/HunyuanDiT

    4,292عرض على GitHub↗

    HunyuanDiT هو نموذج توليدي ثنائي اللغة لتحويل النص إلى صورة ومولد صور يعتمد على محولات الانتشار (diffusion transformer). يستخدم نظام انتشار كامن لتوليف صور عالية الدقة من مطالبات نصية، مع تركيز خاص على فهم وتوليد المحتوى من الأوصاف باللغتين الصينية والإنجليزية. يتميز المشروع ببنية محول متعددة الدقة ومساحة تضمين ثنائية اللغة لربط نصوص مختلفة في منطقة دلالية مشتركة. يدعم تحسين الصور التكراري متعدد الجولات، والذي يترجم الحوار التفاعلي إلى مطالبات محدثة لتعديل المحتوى المرئي تدريجياً. يتضمن النظام قدرات للتعليق التوضيحي التلقائي للصور، وقيود هيكلية للصور للتحكم في التخطيط، وضبط أوزان النموذج لتكييف المولد مع مجموعات بيانات أو أنماط فنية محددة. تشمل تحسينات الأداء تقطير النموذج لتسريع الاستدلال ودعم التنفيذ على الأجهزة ذات ذاكرة الفيديو المنخفضة.

    Implements step-distillation techniques to reduce the number of sampling steps and accelerate image generation.

    Jupyter Notebook
    عرض على GitHub↗4,292
  • hao-ai-lab/fastvideoالصورة الرمزية لـ hao-ai-lab

    hao-ai-lab/FastVideo

    3,743عرض على GitHub↗

    FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine, a video diffusion training framework, and a modular pipeline orchestrator. It provides a distributed transformer optimizer and a distillation toolkit designed to reduce denoising steps and model complexity to increase frame rates. The project distinguishes itself through specialized acceleration techniques, including joint distillation and sparse attention training. It implements low-step video generation and weight quantization to FP8 or FP4 precision to increase throughput a

    Reduces inference latency and denoising steps through distillation and sparse attention for faster video production.

    Pythondiffusersdiffusion-modelsdistillation
    عرض على GitHub↗3,743
  • sandai-org/magi-1الصورة الرمزية لـ SandAI-org

    SandAI-org/MAGI-1

    3,711عرض على GitHub↗

    MAGI-1 is an autoregressive video generation model designed to synthesize high-resolution video sequences from text prompts and image references. It functions as a generative system for text-to-video, image-to-video, and video-to-video transformations. The model utilizes an autoregressive architecture that treats spatio-temporal patches as a sequence of discrete tokens to maintain temporal motion. It employs a variational autoencoder to compress the spatial and temporal dimensions of video data and uses distillation-based step scaling to allow for inference budget control. The system integra

    Utilizes distillation-based step scaling to reduce the number of sampling steps required for high-quality video generation.

    Pythonautoregressivediffusion-modelsvideo-generation
    عرض على GitHub↗3,711
  • thu-ml/turbodiffusionالصورة الرمزية لـ thu-ml

    thu-ml/TurboDiffusion

    3,339عرض على GitHub↗

    TurboDiffusion is a video diffusion inference engine and generator designed to create high-resolution videos from text prompts and images. It provides a runtime environment for executing optimized diffusion model checkpoints with a focus on reducing latency and GPU memory usage. The project features a specialized training framework for aligning sparse-linear attention models with pretrained full-attention models. This system includes capabilities for sparse attention parameter merging and sparse-linear model alignment to reduce computational costs during inference while maintaining output qua

    Reduces generation time and the number of inference steps through attention acceleration and timestep distillation.

    Pythonai-infraconsistency-modeldiffusion-models
    عرض على GitHub↗3,339
  1. Home
  2. Artificial Intelligence & ML
  3. Generative AI Resources
  4. Diffusion & Visual Synthesis Models
  5. Generative AI Models
  6. Diffusion Models
  7. Inference Acceleration

استكشف الوسوم الفرعية

  • Seed Exploration AccelerationTechniques for rapidly sampling candidates to identify high-quality initial seeds using low precision. **Distinct from Inference Acceleration:** Focuses on seed discovery and candidate filtering, while [f14_mt2] is general sampling speed reduction.
  • Step-Distilled AcceleratorsAccelerators that reduce diffusion model sampling steps through distillation, cache inference, and sparse attention. **Distinct from Inference Acceleration:** Distinct from Inference Acceleration: specifically targets step distillation and cache techniques for diffusion models rather than general inference optimization.