awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

50 repositorios

Awesome GitHub RepositoriesSequence Decoding Models

Models that map visual features to text sequences without explicit character segmentation.

Distinguishing note: Specifically addresses sequence-based decoding for OCR, distinct from general sequence-to-sequence models.

Explore 50 awesome GitHub repositories matching artificial intelligence & ml · Sequence Decoding Models. Refine with filters or upvote what's useful.

Awesome Sequence Decoding Models GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • azl397985856/leetcodeAvatar de azl397985856

    azl397985856/leetcode

    55,758Ver en GitHub↗

    This project is a curated educational resource and solution repository for algorithmic challenges, specifically focused on LeetCode problems. It serves as a technical reference for common data structures and algorithmic patterns, providing verified code implementations across multiple programming languages alongside detailed logic and complexity analysis. The repository functions as a comprehensive study guide for competitive programming and technical interview preparation. It includes specialized learning tools such as an Anki flashcard dataset for spaced repetition and a browser extension t

    The project computes the total ways to decode a numeric string into letters using dynamic programming.

    JavaScriptalgoalgorithmalgorithms
    Ver en GitHub↗55,758
  • jaidedai/easyocrAvatar de JaidedAI

    JaidedAI/EasyOCR

    29,615Ver en GitHub↗

    EasyOCR is a deep learning-based computer vision library designed to perform optical character recognition on images and video frames. It functions as a comprehensive pipeline that automates the transformation of visual text into machine-readable strings, enabling the digitization of physical documents, forms, and receipts into searchable data. The engine distinguishes itself through a multi-stage processing workflow that combines convolutional neural networks for spatial feature extraction with sequence-based decoding mechanisms. This architecture allows the system to identify and interpret

    Maps visual character features to text strings using sequence-based decoding mechanisms.

    Pythoncnncrnndata-mining
    Ver en GitHub↗29,615
  • sgl-project/sglangAvatar de sgl-project

    sgl-project/sglang

    29,079Ver en GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Separates compute-intensive prefill and memory-intensive decoding phases across distinct hardware nodes to maximize throughput.

    Pythonattentionblackwellcuda
    Ver en GitHub↗29,079
  • d2l-ai/d2l-enAvatar de d2l-ai

    d2l-ai/d2l-en

    29,001Ver en GitHub↗

    This project is an educational platform and research toolkit designed to teach deep learning through a combination of mathematical theory, visual diagrams, and executable code. It provides a comprehensive environment for building, training, and evaluating neural networks, grounding complex concepts in interactive computational notebooks that allow for hands-on experimentation. The framework distinguishes itself by interleaving theoretical foundations—including linear algebra, calculus, and probability—with practical implementations across multiple industry-standard libraries. It supports flex

    Integrates attention mechanisms into sequence-to-sequence models by dynamically updating context variables.

    Pythonbookcomputer-visiondata-science
    Ver en GitHub↗29,001
  • arendst/tasmotaAvatar de arendst

    arendst/Tasmota

    24,502Ver en GitHub↗

    Tasmota is a universal firmware platform for ESP8266 and ESP32 microcontrollers, designed to provide local control and management of smart home hardware. It functions as an event-driven automation controller that replaces proprietary factory firmware, allowing users to manage relays, sensors, and lighting systems without relying on external cloud services. The system is built on a modular driver architecture that enables dynamic hardware configuration and peripheral support through a web-based management interface. The platform distinguishes itself through a template-driven hardware mapping s

    Translates raw sensor data into structured messages using custom decoder files to simplify integration with local automation systems.

    Carduinoautomationesp32
    Ver en GitHub↗24,502
  • microsoft/unilmAvatar de microsoft

    microsoft/unilm

    22,030Ver en GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Predicts multiple tokens simultaneously during sequence generation to reduce decoding steps.

    Pythonbeitbeit-3bitnet
    Ver en GitHub↗22,030
  • deepseek-ai/flashmlaAvatar de deepseek-ai

    deepseek-ai/FlashMLA

    12,706Ver en GitHub↗

    FlashMLA is an LLM attention kernel library and inference acceleration library providing a collection of high-performance CUDA kernels. It implements multi-head latent attention mechanisms designed to reduce memory overhead and increase throughput during the forward and backward passes of large language model inference. The library utilizes quantized cache attention kernels to improve computation efficiency across both sparse and dense token processing. It specifically optimizes the prefill and decoding phases of model inference through these latent attention implementations. The project cov

    Implements distinct computational paths to optimize the transition between prefill and decoding phases.

    C++
    Ver en GitHub↗12,706
  • sapientinc/hrmAvatar de sapientinc

    sapientinc/HRM

    12,546Ver en GitHub↗

    HRM is an automated reasoning engine and language framework designed to execute complex, multi-scale problem solving. It functions as a reinforcement learning agent that continuously updates internal knowledge representations to improve task performance based on incoming data streams. The system distinguishes itself through a hierarchical architecture that coordinates abstract, long-term planning with granular, low-level logic. By integrating evolutionary algorithms and reinforcement learning, the framework refines model parameters and weights over successive generations, ensuring that intern

    Decodes high-dimensional latent representations into structured natural language sequences.

    Pythonbrain-inspired-aideep-learninglarge-language-models
    Ver en GitHub↗12,546
  • ludwig-ai/ludwigAvatar de ludwig-ai

    ludwig-ai/ludwig

    11,717Ver en GitHub↗

    Ludwig is a multimodal machine learning platform and low-code framework designed for building, training, and deploying neural networks. It enables the construction of models that process text, images, audio, and tabular data through a unified interface using declarative configuration files rather than custom code. The system features a specialized low-code framework for large language models, supporting supervised fine-tuning, preference alignment, and a constrained decoding tool to force structured data output via logit extraction. It also includes an automated model architecture search to i

    Forces large language models to produce structured data using logit extraction and constrained decoding.

    Pythoncomputer-visiondata-centricdata-science
    Ver en GitHub↗11,717
  • idea-research/groundingdinoAvatar de IDEA-Research

    IDEA-Research/GroundingDINO

    9,738Ver en GitHub↗

    GroundingDINO is a deep learning vision model and open-vocabulary object detector designed to map natural language prompts to spatial coordinates. It functions as a text-to-bounding-box framework that enables zero-shot image localization, allowing the system to identify and locate arbitrary objects without requiring predefined classes or specific training for those categories. The project distinguishes itself by matching visual features to natural language descriptions to achieve open-set visual recognition. It supports text-guided image localization and the isolation of specific objects base

    Employs a sequence decoder to predict spatial coordinates and class labels from multimodal features.

    Pythonobject-detectionopen-worldopen-world-detection
    Ver en GitHub↗9,738
  • morvanzhou/pytorch-tutorialAvatar de MorvanZhou

    MorvanZhou/PyTorch-Tutorial

    8,458Ver en GitHub↗

    This project is a collection of PyTorch learning resources and educational guides designed to teach the construction and training of neural networks. It serves as a comprehensive deep learning tutorial covering various model architectures and practical implementation strategies. The resources provide specific guidance on implementing computer vision tasks, such as image classification and synthetic imagery generation, as well as reinforcement learning agents using value networks and experience replay. It also covers sequential data modeling through recurrent networks and generative modeling u

    Provides guidance on implementing architectures that map input sequences to target sequences via latent representations.

    Jupyter Notebookautoencoderbatchbatch-normalization
    Ver en GitHub↗8,458
  • czy36mengfei/tensorflow2_tutorials_chineseAvatar de czy36mengfei

    czy36mengfei/tensorflow2_tutorials_chinese

    7,786Ver en GitHub↗

    This project is a collection of educational resources and instructional guides for learning deep learning and neural network implementation using TensorFlow. It provides a structured set of tutorials and notebooks written in Chinese, covering supervised and unsupervised learning tasks. The material focuses on practical implementations of diverse neural network architectures, including convolutional, recurrent, and autoencoder networks. It includes specific training content for computer vision, natural language processing, and generative models. The coverage extends to specialized network arc

    Implements sequence-to-sequence mappings for tasks like machine translation using attention mechanisms.

    Jupyter Notebook
    Ver en GitHub↗7,786
  • infrasys-ai/aiinfraAvatar de Infrasys-AI

    Infrasys-AI/AIInfra

    7,414Ver en GitHub↗

    Implements separation of prefill and decode phases to avoid resource contention.

    Jupyter Notebookaiinfraaisystem
    Ver en GitHub↗7,414
  • harvardnlp/annotated-transformerAvatar de harvardnlp

    harvardnlp/annotated-transformer

    7,325Ver en GitHub↗

    The Annotated Transformer is an educational resource that provides annotated code implementations of the Transformer architecture for sequence-to-sequence tasks, built with PyTorch. It serves as a learning tool for understanding attention mechanisms, multi-head parallel attention, and scaled dot-product attention through executable examples that walk through each component of the model. The project covers the full Transformer pipeline, including stacked encoder-decoder layers with residual connections and layer normalization, sinusoidal positional encoding for order-aware representation, and

    Generates an output sequence token by token using masked self-attention and encoder-decoder attention.

    Jupyter Notebookannotatednotebookpython
    Ver en GitHub↗7,325
  • princewen/tensorflow_practiceAvatar de princewen

    princewen/tensorflow_practice

    7,009Ver en GitHub↗

    This repository is a collection of practical deep learning implementations and examples built using the TensorFlow framework. It provides a variety of neural network architectures focusing on natural language processing, recommendation systems, reinforcement learning, and time series prediction. The project features a range of specialized models, including sequence-to-sequence and transformer architectures for text processing, and factorization machines for personalized ranking and retrieval. It also includes implementations of reinforcement learning agents using actor-critic and policy gradi

    Implements encoder-decoder architectures that map input sequences to target output sequences.

    Python
    Ver en GitHub↗7,009
  • lmcache/lmcacheAvatar de LMCache

    LMCache/LMCache

    6,909Ver en GitHub↗

    LMCache is a distributed key-value cache manager and tiering system designed to accelerate large language model inference. It functions as a tiered storage layer that offloads tensors from GPU memory to CPU RAM, local disks, or remote object stores, enabling the reuse of cached prefixes across different inference sessions and serving engines. The system differentiates itself through a disaggregated prefill-decode model, which separates prompt processing from token generation by transferring caches between distributed compute nodes. It utilizes peer-to-peer orchestration to share and retrieve

    Implements an architecture that separates prompt processing from token generation by transferring KV caches across compute nodes.

    Pythonamdcudafast
    Ver en GitHub↗6,909
  • albertan017/llm4decompileAvatar de albertan017

    albertan017/LLM4Decompile

    6,728Ver en GitHub↗

    LLM4Decompile es un conjunto de herramientas y framework para la traducción de código binario a código fuente. Utiliza modelos de lenguaje de gran tamaño (LLM) para transformar código máquina en código fuente legible y recuperar la lógica original de ejecutables compilados. El proyecto incluye un pipeline especializado para generar datasets de entrenamiento sintéticos convirtiendo código fuente en pares de ensamblador. Proporciona un framework de fine-tuning para optimizar modelos de deep learning en estos datasets de binario a fuente, aumentando la precisión de la recuperación de código. El sistema también cuenta con capacidades para refinar pseudocódigo descompilado. Este proceso se centra en restaurar el esqueleto estructural y los nombres de variables de un binario para mejorar la legibilidad de la lógica desensamblada.

    Implements neural machine translation architectures to map binary disassembly sequences to high-level code.

    Python
    Ver en GitHub↗6,728
  • ericlbuehler/mistral.rsAvatar de EricLBuehler

    EricLBuehler/mistral.rs

    6,597Ver en GitHub↗

    mistral.rs is an inference engine for large language models that runs locally and exposes models behind OpenAI and Anthropic-compatible APIs. It serves as a multi-model serving platform, capable of loading several models in a single server process with per-request routing and on-demand loading and unloading. The engine supports multimodal inference, processing text alongside images, video, audio, and speech inputs, and includes a quantized model deployment runtime that reduces memory use and speeds up inference on consumer hardware. The project distinguishes itself through an agentic tool exe

    Enforces JSON Schema on tool call arguments during decoding to prevent malformed output.

    Rustllmrustuqff
    Ver en GitHub↗6,597
  • tensorflow/nmtAvatar de tensorflow

    tensorflow/nmt

    6,461Ver en GitHub↗

    This project is a neural machine translation system used to build models that automatically translate text from one language to another. It utilizes sequence-to-sequence modeling to transform variable-length input sequences into corresponding output sequences. The system implements bidirectional recurrent neural network encoding and attention mechanisms to capture contextual information and focus on specific parts of the source text during translation. To manage training and inference, it employs separate computational graphs and supports distributing model layers across multiple GPU devices.

    Implements a sequence-to-sequence architecture to map source text to target translation via a latent representation.

    Python
    Ver en GitHub↗6,461
  • facebookresearch/wav2letterAvatar de facebookresearch

    facebookresearch/wav2letter

    6,444Ver en GitHub↗

    wav2letter is an automatic speech recognition toolkit and deep learning framework designed to convert audio speech signals into written text. It functions as a distributed training system and an inference engine for building and deploying neural network architectures. The system enables the training of large-scale speech models across multiple compute nodes using custom architecture files and structured recipes. It includes an inference engine that allows these trained models to be executed within Python workflows to transform audio sequences into text. The framework covers the full speech r

    Implements sequence decoders to determine the most accurate sequence of words for a given audio input.

    C++
    Ver en GitHub↗6,444
Ant.123Siguiente
  1. Home
  2. Artificial Intelligence & ML
  3. Sequence Decoding Models

Explorar subetiquetas

  • Sequence Decoders8 sub-etiquetasComponents for generating output sequences conditioned on input context. **Distinct from Sequence Decoding Models:** Distinct from general sequence decoding models: focuses on the decoder component logic rather than the full model architecture.
  • Sequence-to-Sequence Mappings1 sub-etiquetaArchitectures that map an input sequence to a target sequence via a latent representation. **Distinct from Sequence Decoding Models:** Broadens from OCR-specific sequence decoding to general sequence-to-sequence mapping including translation.