awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 रिपॉजिटरी

Awesome GitHub RepositoriesGeneration Speed Optimizers

Utilities that improve generation latency by reducing the number of model forward passes.

Distinct from Token Optimization Utilities: Distinct from Token Optimization Utilities: targets GPU pass reduction for speed rather than just token count reduction for cost.

Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Generation Speed Optimizers. Refine with filters or upvote what's useful.

Awesome Generation Speed Optimizers GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • microsoft/guidancemicrosoft का अवतार

    microsoft/guidance

    21,502GitHub पर देखें↗

    Guidance is a control framework and generation orchestrator for large language models. It provides a programming layer to steer model outputs through structured templates, schema enforcement, and logical flow management. The framework distinguishes itself by interleaving model generation with local code execution, enabling the use of loops and conditional branching within a single session. It employs grammar-based token constraints and regular expressions to force models to sample only from tokens that satisfy a specific structural format, ensuring strict adherence to predefined data models.

    Improves generation speed by inserting known tokens directly into the output stream to reduce GPU passes.

    Jupyter Notebook
    GitHub पर देखें↗21,502
  • yujiangshui/a-programmers-guide-to-englishyujiangshui का अवतार

    yujiangshui/A-Programmers-Guide-to-English

    16,428GitHub पर देखें↗

    This project is a systematic framework for English language acquisition that applies structured workflows and cognitive strategies to build linguistic proficiency. It focuses on the construction of a linguistic knowledge base, enabling learners to master vocabulary and grammar through methodical training. The methodology is distinguished by its use of computer science concepts, such as mental-model-based learning and memory buffers, to organize progression. It emphasizes a cognitive-translation bypass to develop target language thinking, reducing mental latency by processing information direc

    Removes the need for mental translation to improve real-time interaction efficiency.

    englishenglish-learning
    GitHub पर देखें↗16,428
  • cumulo-autumn/streamdiffusioncumulo-autumn का अवतार

    cumulo-autumn/StreamDiffusion

    10,770GitHub पर देखें↗

    StreamDiffusion is an interactive generative AI framework and inference engine designed for the low-latency delivery of image and video streams. It provides a real-time Stable Diffusion pipeline for text-to-image and image-to-image generation, enabling the creation of continuous generative image streams with minimized computational delay. The framework optimizes throughput using a pre-computed cache engine and residual-based guidance approximation to reduce the number of required model passes. It further manages GPU load through similarity-based frame skipping, which avoids redundant computat

    Accelerates image generation by reducing the number of required model forward passes.

    Python
    GitHub पर देखें↗10,770
  • openai/consistency_modelsopenai का अवतार

    openai/consistency_models

    6,492GitHub पर देखें↗

    This project is a framework for training and sampling generative models designed to produce high-quality images in few steps. It provides implementations for image generation models that transform random noise into structured visual data through an optimized sampling process. The system specializes in accelerating image generation through consistency distillation and consistency training. It includes tools to transform pre-trained diffusion models into faster versions by distilling knowledge from a teacher model into a student model, as well as methods to train consistency models from scratch

    Optimizes inference speed by reducing the number of sampling steps required for image generation.

    Python
    GitHub पर देखें↗6,492
  • getstream/vision-agentsGetStream का अवतार

    GetStream/Vision-Agents

    6,029GitHub पर देखें↗

    Chooses model size and compute precision to balance transcription accuracy against inference speed on available hardware.

    Pythonagentic-aiagentsai
    GitHub पर देखें↗6,029
  • xiaomi/maceXiaoMi का अवतार

    XiaoMi/mace

    5,041GitHub पर देखें↗

    Mace एक मोबाइल डीप लर्निंग इन्फरेंस फ्रेमवर्क और हार्डवेयर एक्सेलेरेशन इंजन है। यह मोबाइल उपकरणों पर न्यूरल नेटवर्क मॉडल निष्पादित करने के लिए एक रनटाइम के रूप में कार्य करता है, जो CPU, GPU और NPU पर गणनाओं को वितरित करता है। इस प्रोजेक्ट में विभिन्न इंडस्ट्री फॉर्मेट से प्री-ट्रेन किए गए न्यूरल नेटवर्क को मोबाइल-ऑप्टिमाइज़्ड रिप्रेजेंटेशन में बदलने के लिए एक क्रॉस-प्लेटफ़ॉर्म मॉडल कन्वर्टर शामिल है। यह एक न्यूरल नेटवर्क ऑब्फ्यूस्केटर भी प्रदान करता है जो रिवर्स इंजीनियरिंग से बौद्धिक संपदा की रक्षा के लिए मॉडल वेट्स को सोर्स कोड में बदल देता है। फ्रेमवर्क मेमोरी आवंटन को ऑप्टिमाइज़ करके और चिप पावर सेटिंग्स को समायोजित करके ऑन-डिवाइस संसाधनों को मैनेज करता है।

    Increases operation speed by applying hardware acceleration and optimized mathematical algorithms to complex calculations.

    C++deep-learninghvxmachine-learning
    GitHub पर देखें↗5,041
  • cvg/lightgluecvg का अवतार

    cvg/LightGlue

    4,625GitHub पर देखें↗

    LightGlue एक डीप लर्निंग फ्रेमवर्क है जिसे इमेजेस के जोड़ों के बीच लोकल फीचर मैचिंग और हाई-स्पीड कॉरेस्पोंडेंस एस्टिमेशन के लिए डिज़ाइन किया गया है। यह एक कंप्यूटर विज़न मैचिंग मॉडल के रूप में कार्य करता है जो अलग-अलग दृष्टिकोणों (viewpoints) में संबंधित की-पॉइंट्स की पहचान करता है। यह सिस्टम एक एडेप्टिव न्यूरल नेटवर्क आर्किटेक्चर का उपयोग करता है जो इनपुट इमेज पेयर्स के आधार पर अपनी गहराई और चौड़ाई को प्रून (prune) करके इन्फरेंस स्पीड को गतिशील रूप से ऑप्टिमाइज़ करता है। यह दृष्टिकोण फीचर डिस्क्रिप्टर्स के बीच सहसंबंधों (correlations) की गणना करने के लिए ट्रांसफॉर्मर-शैली के अटेंशन मैकेनिज्म और क्रॉस-इमेज अटेंशन का उपयोग करता है। मैचिंग प्रक्रिया में एक इटरेटिव रिफाइनमेंट लूप और डायनामिक अर्ली स्टॉपिंग शामिल है ताकि कॉन्फिडेंस थ्रेशोल्ड पूरा होने पर गणना को रोका जा सके। ये क्षमताएं रीयल-टाइम इमेज अलाइनमेंट और न्यूरल नेटवर्क इन्फरेंस ऑप्टिमाइज़ेशन के लिए एक व्यापक कंप्यूटर विज़न पाइपलाइन का समर्थन करती हैं।

    Reduces computational cost and increases processing speed through adaptive network pruning during inference.

    Python
    GitHub पर देखें↗4,625
  • tingsongyu/pytorch-tutorial-2ndTingsongYu का अवतार

    TingsongYu/PyTorch-Tutorial-2nd

    4,555GitHub पर देखें↗

    This project is a comprehensive instructional resource and course for building neural networks using PyTorch. It covers the fundamental building blocks of deep learning, including tensor manipulation, automatic differentiation, and the construction of modular neural network components. The repository serves as a technical guide for several specialized domains. It provides implementation details for computer vision tasks such as image classification, object detection, and semantic segmentation, as well as natural language processing workflows involving transformers, recurrent networks, and gen

    Adjusts image size and confidence thresholds to balance execution speed and accuracy.

    Jupyter Notebookcomputer-visiondeepsortdiffusion-models
    GitHub पर देखें↗4,555
  • huawei-noah/ghostnethuawei-noah का अवतार

    huawei-noah/ghostnet

    4,416GitHub पर देखें↗

    GhostNet कुशल AI मॉडल आर्किटेक्चर और न्यूरल नेटवर्क डिज़ाइन पैटर्न का एक सेट प्रदान करता है, जिसे गणना और मेमोरी ओवरहेड को कम करने के लिए डिज़ाइन किया गया है। यह एक कंप्यूटर विज़न बैकबोन और हल्के विज़न ट्रांसफार्मर के रूप में कार्य करता है, जो पूर्वानुमानित सटीकता और इन्फरेंस गति के बीच संतुलन को अनुकूलित करता है। यह प्रोजेक्ट मोबाइल उपकरणों और एज हार्डवेयर पर तैनाती के लिए संसाधन खपत को कम करने पर केंद्रित है। यह हल्के विज़न ट्रांसफार्मर कार्यान्वयन और ऐसे आर्किटेक्चर के उपयोग के माध्यम से इसे प्राप्त करता है जो पैरामीटर की कुल संख्या को कम करते हैं। कोडबेस इन्फरेंस ऑप्टिमाइज़ेशन के लिए कई क्षमताओं को कवर करता है, जिसमें कम्प्यूटेशनल लागत और मेमोरी उपयोग में कमी शामिल है। यह इन्फरेंस लेटेंसी को कम करने के लिए डेप्थवाइज़-सेपरेबल कन्वेन्शनल ब्लॉक्स, लीनियर-बॉटलनेक डेप्थवाइज़ कन्वेन्शन्स और टाइड-वेट ट्रांसफार्मर ब्लॉक्स जैसे स्ट्रक्चरल डिज़ाइन पैटर्न लागू करता है।

    Optimizes the execution speed and memory usage of neural network inference for faster live predictions.

    Python
    GitHub पर देखें↗4,416
  • apachecn/pytorch-doc-zhapachecn का अवतार

    apachecn/pytorch-doc-zh

    4,224GitHub पर देखें↗

    This project is a Chinese language translation of the technical guides and API references for the PyTorch deep learning framework. It serves as a localized knowledge base and reference material to make deep learning documentation accessible to non-English speakers. The documentation covers a comprehensive range of PyTorch capabilities, including neural network model development, automatic differentiation, and the implementation of backend kernels. It provides detailed guidance on distributed training strategies, model deployment through formats like ONNX and C++, and various model optimizatio

    Offers methods to reduce memory footprint and execution time by converting models to quantized formats.

    Shelldeep-learningdocumentationpython
    GitHub पर देखें↗4,224
  • metavoiceio/metavoice-srcmetavoiceio का अवतार

    metavoiceio/metavoice-src

    4,202GitHub पर देखें↗

    This project is an expressive text-to-speech foundation model and voice cloning system designed to synthesize human-like speech with emotional nuance and high fidelity. It functions as a finetunable speech model that can generate audio mimicking a specific person using a reference voice sample. The system distinguishes itself through a high-performance inference engine that utilizes memory caching and hardware compilation to reduce latency during the audio generation process. It further allows for synthesis quality improvements by training the language model on custom datasets consisting of a

    Optimizes the execution speed of neural network inference via hardware compilation and memory caching.

    Pythonaideep-learningpytorch
    GitHub पर देखें↗4,202
  • vectorspacelab/omnigen2VectorSpaceLab का अवतार

    VectorSpaceLab/OmniGen2

    4,093GitHub पर देखें↗

    OmniGen2 एक एकीकृत इमेज जनरेशन मॉडल और मल्टीमॉडल लार्ज लैंग्वेज मॉडल है जिसे एक ही फ्रेमवर्क के भीतर टेक्स्ट-टू-इमेज जनरेशन, इमेज-टू-इमेज कार्यों और इमेज एडिटिंग को संभालने के लिए डिज़ाइन किया गया है। यह एक कॉज़ल लैंग्वेज मॉडल विज़ुअल इंजन के रूप में कार्य करता है जो संयुक्त टेक्स्ट और विज़ुअल इनपुट के आधार पर इमेज उत्पन्न और संपादित करने में सक्षम है। यह सिस्टम इन-कॉन्टेक्स्ट विज़ुअल कंपोज़िशन और सब्जेक्ट-ड्रिवन जनरेशन की सुविधा देता है, जिससे यह रेफरेंस इमेज से विषयों को निकालने और उन्हें नए दृश्यों में रखने की अनुमति देता है। यह निर्देश-आधारित इमेज एडिटिंग का भी समर्थन करता है, जहां विशिष्ट वस्तुओं या शैलियों को प्राकृतिक भाषा कमांड के माध्यम से संशोधित किया जाता है जबकि इमेज का बाकी हिस्सा संरक्षित रहता है। मॉडल की क्षमताएं विज़ुअल कंटेंट विश्लेषण और तर्क तक फैली हुई हैं, जो संयुक्त टेक्स्ट और विज़न इनपुट में वस्तुओं की पहचान को सक्षम बनाती हैं। आउटपुट गुणवत्ता में सुधार करने के लिए, यह सेल्फ-करेक्शन तंत्र के साथ एक पुनरावृत्ति विज़ुअल रिफाइनमेंट प्रक्रिया का उपयोग करता है। प्रदर्शन को डायनामिक वेट ऑफलोडिंग के माध्यम से VRAM उपयोग ऑप्टिमाइज़ेशन और कैशिंग तकनीकों का उपयोग करके इन्फरेंस गति त्वरण के माध्यम से प्रबंधित किया जाता है।

    Increases generation throughput using caching techniques and adjusted guidance ranges to accelerate model inference.

    Jupyter Notebook
    GitHub पर देखें↗4,093
  • sandai-org/magi-1SandAI-org का अवतार

    SandAI-org/MAGI-1

    3,711GitHub पर देखें↗

    MAGI-1 is an autoregressive video generation model designed to synthesize high-resolution video sequences from text prompts and image references. It functions as a generative system for text-to-video, image-to-video, and video-to-video transformations. The model utilizes an autoregressive architecture that treats spatio-temporal patches as a sequence of discrete tokens to maintain temporal motion. It employs a variational autoencoder to compress the spatial and temporal dimensions of video data and uses distillation-based step scaling to allow for inference budget control. The system integra

    Provides inference budget control to balance output quality and processing speed via distillation-based step sizes.

    Pythonautoregressivediffusion-modelsvideo-generation
    GitHub पर देखें↗3,711
  • hkuds/paper2slidesHKUDS का अवतार

    HKUDS/Paper2Slides

    3,092GitHub पर देखें↗

    Paper2Slides is an AI-driven presentation generator and content extractor designed to transform academic papers and scientific documents into structured slides and posters. It utilizes retrieval-augmented generation to distill key data points and identify critical figures while maintaining direct traceability to the original source text. The system functions as an AI slide designer that applies professional themes or custom visual styles defined through natural language. It integrates with external image generation services to produce high-quality visuals and research visualizations for acade

    Optimizes processing speed by allowing a choice between deep semantic indexing for complex papers and a fast mode for short files.

    Pythonagentic-aillm-agentspaper2poster
    GitHub पर देखें↗3,092
  1. Home
  2. Artificial Intelligence & ML
  3. Token Optimization Utilities
  4. Generation Speed Optimizers

सब-टैग एक्सप्लोर करें

  • Adaptive IndexingOptimization techniques that switch between indexing depths based on document complexity to balance speed and accuracy. **Distinct from Generation Speed Optimizers:** Focuses on switching indexing strategies for document processing rather than reducing GPU passes for LLM inference
  • Cognitive Processing SpeedStrategies to reduce mental translation time for more efficient real-time communication. **Distinct from Generation Speed Optimizers:** Distinct from Generation Speed Optimizers: focuses on human cognitive latency rather than AI model inference speed.
  • Transcription Speed Optimizers1 सब-टैगUtilities that choose model size and compute precision to balance transcription accuracy against inference speed on available hardware. **Distinct from Generation Speed Optimizers:** Distinct from Generation Speed Optimizers: focuses on transcription-specific speed/accuracy tradeoffs, not general generation latency.