awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

20 个仓库

Awesome GitHub RepositoriesTokenizers

Components that decompose raw text into sub-word units or tokens based on statistical frequency and normalization rules.

Explore 20 awesome GitHub repositories matching artificial intelligence & ml · Tokenizers. Refine with filters or upvote what's useful.

Awesome Tokenizers GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • openai/whisperopenai 的头像

    openai/whisper

    102,828在 GitHub 上查看↗

    This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies

    Converts raw text into subword units using byte-level sequences to handle diverse languages without requiring language-specific rules.

    Python
    在 GitHub 上查看↗102,828
  • openai/codexopenai 的头像

    openai/codex

    91,445在 GitHub 上查看↗

    Codex is an automated programming tool and generative code assistant designed to interpret developer intent through a natural language interface. It functions as a machine learning model trained on public code repositories to provide intelligent code completion, suggestions, and refactoring within development environments. By translating human instructions into executable code snippets, the system bridges the gap between high-level technical requirements and functional software implementation. The engine utilizes transformer-based sequence modeling and supervised fine-tuning to align its outp

    Decomposes raw text into sub-word units to represent diverse programming languages and syntax structures efficiently.

    Rust
    在 GitHub 上查看↗91,445
  • karpathy/nanogptkarpathy 的头像

    karpathy/nanoGPT

    59,730在 GitHub 上查看↗

    nanoGPT is a lightweight engine for training and fine-tuning transformer-based language models from scratch. It provides a minimalist codebase designed for educational exploration and rapid experimentation with neural network architectures, utilizing self-attention and feed-forward layers to process sequences and predict subsequent elements. The project distinguishes itself through a focus on high-speed data ingestion and hardware-accelerated performance. It includes a dedicated pipeline for transforming raw text into memory-mapped binary files, which enables efficient streaming during traini

    Decomposes raw text into numerical units using character-level tokenization for model ingestion.

    Python
    在 GitHub 上查看↗59,730
  • datawhalechina/hello-agentsdatawhalechina 的头像

    datawhalechina/hello-agents

    59,685在 GitHub 上查看↗

    This project provides a comprehensive framework for building, training, and managing autonomous agents. It enables the construction of systems that utilize language models to plan, manage memory, and execute multi-step tasks through iterative reasoning loops and tool-based actions. The framework distinguishes itself by offering specialized capabilities for interacting with graphical user interfaces and legacy software, allowing agents to perceive visual elements and perform actions like a human user. It supports complex, cross-application workflows through graph-based orchestration and provid

    Converts raw text into sub-word units using frequency-based algorithms to create efficient vocabulary representations.

    Pythonagentllmrag
    在 GitHub 上查看↗59,685
  • meta-llama/llamameta-llama 的头像

    meta-llama/llama

    59,464在 GitHub 上查看↗

    Llama is a computational framework and runtime environment designed for executing transformer-based neural networks locally. It functions as a generative AI inference engine, enabling the processing of input sequences through pre-trained model weights to produce text completions and structured data outputs directly on your own hardware. The system distinguishes itself through specialized memory and computation management techniques, including memory-mapped weight loading and quantization-aware inference, which allow for efficient execution on standard consumer hardware. It utilizes a stateles

    Decomposes raw text into numerical tokens suitable for processing by transformer-based neural networks.

    Python
    在 GitHub 上查看↗59,464
  • meilisearch/meilisearchmeilisearch 的头像

    meilisearch/meilisearch

    58,118在 GitHub 上查看↗

    Meilisearch is a Rust-based search engine providing typo-tolerant full-text and vector-based semantic search with real-time conversational capabilities.

    Decomposes unstructured text into normalized tokens using language-specific rules to prepare data for indexing.

    Rustaiapiapp-search
    在 GitHub 上查看↗58,118
  • explosion/spacyexplosion 的头像

    explosion/spaCy

    33,688在 GitHub 上查看↗

    spaCy is a Python natural language processing framework designed for industrial-scale text processing. It converts raw text into structured data for machine learning pipelines through a combination of statistical language model trainers, transformer-based text processors, and syntactic dependency parsers. The project enables the integration of pretrained transformer architectures to perform complex linguistic analysis and multi-task learning. It also provides a specialized system for neural named entity recognition to identify and categorize key entities within text. The framework covers a b

    Uses linguistic patterns and regular expressions to decompose raw text into discrete tokens.

    Pythonaiartificial-intelligencecython
    在 GitHub 上查看↗33,688
  • d2l-ai/d2l-end2l-ai 的头像

    d2l-ai/d2l-en

    29,001在 GitHub 上查看↗

    This project is an educational platform and research toolkit designed to teach deep learning through a combination of mathematical theory, visual diagrams, and executable code. It provides a comprehensive environment for building, training, and evaluating neural networks, grounding complex concepts in interactive computational notebooks that allow for hands-on experimentation. The framework distinguishes itself by interleaving theoretical foundations—including linear algebra, calculus, and probability—with practical implementations across multiple industry-standard libraries. It supports flex

    Assigns categorical labels to individual tokens in a sequence using shared classification layers.

    Pythonbookcomputer-visiondata-science
    在 GitHub 上查看↗29,001
  • guidance-ai/guidanceguidance-ai 的头像

    guidance-ai/guidance

    21,502在 GitHub 上查看↗

    Guidance is a generative AI orchestration framework designed to manage complex interactions with language models by embedding programmatic control directly into the prompt generation process. It functions as a prompt programming environment that allows developers to interleave raw text with executable logic, enabling the construction of sophisticated, multi-step agentic workflows. The framework distinguishes itself through grammar-constrained token sampling and stateful stream interception, which restrict the model's output distribution based on formal language rules. By enforcing these const

    Restricts language model token generation based on formal language rules to ensure strict adherence to output schemas.

    Jupyter Notebook
    在 GitHub 上查看↗21,502
  • valeriansaliou/sonicvaleriansaliou 的头像

    valeriansaliou/sonic

    21,249在 GitHub 上查看↗

    Sonic is a high-performance, lightweight search backend designed to provide real-time full-text search and autocomplete capabilities for applications. It functions as a persistent indexing server that maps text terms to object identifiers, allowing developers to integrate rapid search functionality without storing raw document content directly within the search engine. The system distinguishes itself through a specialized graph-based index that enables real-time word prediction and typo correction. Communication is handled via a custom, low-latency binary protocol over raw TCP sockets, which

    Processes raw input through language-specific pipelines that perform tokenization, stop-word removal, and diacritic folding.

    Rustbackenddatabasegraph
    在 GitHub 上查看↗21,249
  • flairnlp/flairflairNLP 的头像

    flairNLP/flair

    14,378在 GitHub 上查看↗

    Flair is a transformer-based natural language processing framework used to build and train models for text classification and sequence tagging. It provides a specialized library for generating contextual text embeddings and performing linguistic analysis. The framework includes dedicated tools for named entity recognition, including the identification of specialized biomedical entities across multiple languages. It further supports entity linking to map identified text mentions to unique entries within general or biomedical knowledge bases. The project covers a broad range of language analys

    Provides neural network layers for assigning categorical labels to individual tokens within a text sequence.

    Python
    在 GitHub 上查看↗14,378
  • outlines-dev/outlinesoutlines-dev 的头像

    outlines-dev/outlines

    13,965在 GitHub 上查看↗

    Outlines is a guided text generation framework and structured output engine for large language models. It enforces precise structural constraints on model output during the sampling process to ensure the generation of valid data. The framework ensures that model outputs strictly adhere to predefined data models, including JSON schemas, regular expressions, and formal grammars. This enables the conversion of natural language inputs into structured arguments for function calling and the generation of valid JSON for downstream processing. The system manages model orchestration through prompt te

    Modifies the token probability distribution to ensure generated text adheres to specific regular expressions or grammars.

    Python
    在 GitHub 上查看↗13,965
  • dottxt-ai/outlinesdottxt-ai 的头像

    dottxt-ai/outlines

    13,446在 GitHub 上查看↗

    Outlines is a library designed to ensure machine-readable output from generative models by applying programmatic constraints during the token sampling process. It functions as a toolkit for forcing large language models to generate text that strictly adheres to JSON schemas, regular expressions, and formal grammars, enabling the integration of model responses into existing software systems. The library distinguishes itself by integrating formal language rules directly into the sampling loop. It achieves this by converting regular expressions into deterministic finite automata and utilizing lo

    Restricts the model's next-token probability distribution by zeroing out tokens that violate defined grammar or schema constraints.

    Pythoncfggenerative-aijson
    在 GitHub 上查看↗13,446
  • google/sentencepiecegoogle 的头像

    google/sentencepiece

    11,657在 GitHub 上查看↗

    SentencePiece is a text segmentation engine and tokenization library designed for machine learning workflows. It provides a comprehensive toolkit for transforming raw text into subword units or numerical identifiers, enabling consistent data representation for neural network training and inference. The library supports the training of segmentation models from raw text, allowing for the creation of custom vocabularies tailored to specific domain requirements. The project distinguishes itself through its byte-level encoding and fallback mechanisms, which ensure that every input can be represent

    Decomposes unknown characters into UTF-8 byte sequences to ensure full vocabulary coverage without unknown tokens.

    C++natural-language-processingneural-machine-translationword-segmentation
    在 GitHub 上查看↗11,657
  • karpathy/minbpekarpathy 的头像

    karpathy/minbpe

    10,582在 GitHub 上查看↗

    Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.

    Provides a clean implementation of the BPE training algorithm to learn merge rules from text corpora.

    Python
    在 GitHub 上查看↗10,582
  • microsoft/llmlinguamicrosoft 的头像

    microsoft/LLMLingua

    5,844在 GitHub 上查看↗

    LLMLingua is a prompt compression tool that reduces token count in prompts before they are sent to a large language model, cutting API costs and latency while preserving task performance. It operates as an extractive pipeline using a BERT-level Transformer encoder to classify each token for removal based on full bidirectional context from the prompt, retaining only key information and discarding non-essential tokens. The tool is trained through a knowledge distillation process, where a compact compression model learns from an extractive dataset derived from a large language model's output to

    Removes redundant tokens identified by a small language model to cut API costs and latency.

    Python
    在 GitHub 上查看↗5,844
  • biolab/orange3biolab 的头像

    biolab/orange3

    5,635在 GitHub 上查看↗

    Orange3 is a visual data mining platform that provides an interactive canvas for building data analysis workflows without writing code. At its core, it offers a widget-based visual programming environment where users connect configurable components to perform data preprocessing, machine learning model training, statistical evaluation, and interactive visualization. The platform is built on NumPy-backed data tables with domain descriptors that define variable names, types, and roles, and includes a lazy SQL query proxy for working with database tables without loading all data into memory. The

    Provides a widget to drop constant attributes and unused categorical values from datasets.

    Python
    在 GitHub 上查看↗5,635
  • onnxsim/onnxsimonnxsim 的头像

    onnxsim/onnxsim

    4,353在 GitHub 上查看↗

    onnxsim 是一个深度学习图优化器和模型简化器,旨在降低 ONNX 计算图的复杂度。它充当模型压缩器,用简化的常量输出替换复杂的算子序列,以减少操作开销。 该项目通过常量折叠推理实现简化,将常量算子的子图替换为预计算的常量张量。它利用基于模式的图重写和静态计算图分析来识别并删除冗余节点或不可达操作。 该工具涵盖了广泛的模型优化功能,包括算子冗余消除以及删除不必要的 reshape 或 identity 节点。这些过程简化了执行流程并减少了模型的内存占用。

    Eliminates identity operations and unnecessary reshape nodes that do not alter mathematical output.

    C++deep-learningonnxpytorch
    在 GitHub 上查看↗4,353
  • mbloch/mapshapermbloch 的头像

    mbloch/mapshaper

    4,133在 GitHub 上查看↗

    Mapshaper 是一个用于处理、简化和转换地理矢量数据的工具,提供命令行界面、Web 浏览器工具和 Node.js 库。它作为一个坐标投影器、矢量数据转换器和 Web 地图资产优化器,旨在在不同的坐标参考系统和文件格式之间转换空间数据集。 该项目以其拓扑保持几何简化而著称,在减少顶点数量的同时保持共享边界,以防止间隙和重叠。它还通过坐标量化和属性过滤进一步优化 Web 资产,以减小文件大小。 该系统涵盖了广泛的功能,包括使用 PROJ 字符串和 EPSG 代码进行坐标重投影,以及跨 Shapefile、GeoJSON、TopoJSON、GeoPackage 和 KML 等格式的数据转换。它提供了广泛的几何处理工具,用于缓冲、裁剪、溶解和修复拓扑,以及用于属性连接、过滤和转换的数据管理实用程序。此外,它还包括用于生成样式化 SVG 导出、经纬网和比例符号地图的视觉功能。 空间处理功能可以通过其 Node.js 库直接集成到 JavaScript 应用程序和构建流水线中。

    Deletes features that share the same identifier as a previous feature to clean datasets.

    JavaScript
    在 GitHub 上查看↗4,133
  • huawei-noah/pretrained-language-modelhuawei-noah 的头像

    huawei-noah/Pretrained-Language-Model

    3,163在 GitHub 上查看↗

    Pretrained-Language-Model is a machine learning library and natural language processing toolkit designed for pretraining, tokenizing, and compressing large language models using transformer architectures and specialized optimization techniques. It supports Chinese and multilingual natural language processing tasks, including text classification and conversational response generation. The framework provides specialized capabilities for training large-scale autoregressive and contextual language models, alongside model compression techniques like knowledge distillation and quantization to reduc

    Splits raw text streams into subword tokens using byte-level vocabularies for downstream NLP processing.

    Pythonknowledge-distillationlarge-scale-distributedmodel-compression
    在 GitHub 上查看↗3,163
  1. Home
  2. Artificial Intelligence & ML
  3. Natural Language Processing
  4. Tokenizers

探索子标签

  • Byte-Level Tokenizers2 个子标签Tokenization methods that operate on raw byte sequences to handle diverse vocabularies.
  • Grammar-Constrained Token SamplersTools that restrict token generation based on formal language rules. **Distinct from Tokenizers:** Focuses on grammar-based sampling tools, distinct from general tokenizers.
  • Token Tagging LayersNeural network layers for assigning categorical labels to individual tokens in a sequence. **Distinct from Tokenizers:** Focuses on the tagging layer architecture, distinct from general tokenization.