awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

8 repository-uri

Awesome GitHub RepositoriesByte-Level Tokenizers

Tokenization methods that operate on raw byte sequences to handle diverse vocabularies.

Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Byte-Level Tokenizers. Refine with filters or upvote what's useful.

Awesome Byte-Level Tokenizers GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • openai/whisperAvatar openai

    openai/whisper

    102,828Vezi pe GitHub↗

    This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies

    Converts raw text into subword units using byte-level sequences to handle diverse languages without requiring language-specific rules.

    Python
    Vezi pe GitHub↗102,828
  • google/sentencepieceAvatar google

    google/sentencepiece

    11,657Vezi pe GitHub↗

    SentencePiece is a text segmentation engine and tokenization library designed for machine learning workflows. It provides a comprehensive toolkit for transforming raw text into subword units or numerical identifiers, enabling consistent data representation for neural network training and inference. The library supports the training of segmentation models from raw text, allowing for the creation of custom vocabularies tailored to specific domain requirements. The project distinguishes itself through its byte-level encoding and fallback mechanisms, which ensure that every input can be represent

    Decomposes unknown characters into UTF-8 byte sequences to ensure full vocabulary coverage without unknown tokens.

    C++natural-language-processingneural-machine-translationword-segmentation
    Vezi pe GitHub↗11,657
  • karpathy/minbpeAvatar karpathy

    karpathy/minbpe

    10,582Vezi pe GitHub↗

    Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.

    Provides a clean implementation of the BPE training algorithm to learn merge rules from text corpora.

    Python
    Vezi pe GitHub↗10,582
  • microsoft/llmlinguaAvatar microsoft

    microsoft/LLMLingua

    5,844Vezi pe GitHub↗

    LLMLingua is a prompt compression tool that reduces token count in prompts before they are sent to a large language model, cutting API costs and latency while preserving task performance. It operates as an extractive pipeline using a BERT-level Transformer encoder to classify each token for removal based on full bidirectional context from the prompt, retaining only key information and discarding non-essential tokens. The tool is trained through a knowledge distillation process, where a compact compression model learns from an extractive dataset derived from a large language model's output to

    Removes redundant tokens identified by a small language model to cut API costs and latency.

    Python
    Vezi pe GitHub↗5,844
  • biolab/orange3Avatar biolab

    biolab/orange3

    5,635Vezi pe GitHub↗

    Orange3 is a visual data mining platform that provides an interactive canvas for building data analysis workflows without writing code. At its core, it offers a widget-based visual programming environment where users connect configurable components to perform data preprocessing, machine learning model training, statistical evaluation, and interactive visualization. The platform is built on NumPy-backed data tables with domain descriptors that define variable names, types, and roles, and includes a lazy SQL query proxy for working with database tables without loading all data into memory. The

    Provides a widget to drop constant attributes and unused categorical values from datasets.

    Python
    Vezi pe GitHub↗5,635
  • onnxsim/onnxsimAvatar onnxsim

    onnxsim/onnxsim

    4,353Vezi pe GitHub↗

    onnxsim este un optimizator de grafuri de deep learning și un simplificator de modele conceput pentru a reduce complexitatea grafurilor de calcul ONNX. Funcționează ca un compresor de modele care înlocuiește secvențele complexe de operatori cu ieșiri constante simplificate pentru a reduce overhead-ul operațional. Proiectul obține simplificarea prin inferența constant folding, care înlocuiește subgrafurile de operatori constanți cu tensori constanți pre-calculați. Utilizează rescrierea grafurilor bazată pe modele și analiza statică a grafurilor de calcul pentru a identifica și elimina nodurile redundante sau operațiunile inaccesibile. Instrumentul acoperă capabilități largi de optimizare a modelelor, inclusiv eliminarea redundanței operatorilor și eliminarea nodurilor inutile de reshape sau identity. Aceste procese eficientizează fluxul de execuție și reduc amprenta de memorie a modelului.

    Eliminates identity operations and unnecessary reshape nodes that do not alter mathematical output.

    C++deep-learningonnxpytorch
    Vezi pe GitHub↗4,353
  • mbloch/mapshaperAvatar mbloch

    mbloch/mapshaper

    4,133Vezi pe GitHub↗

    Mapshaper este un instrument pentru procesarea, simplificarea și convertirea datelor vectoriale geografice, disponibil ca interfață de linie de comandă, instrument de browser web și bibliotecă Node.js. Funcționează ca un proiector de coordonate, convertor de date vectoriale și optimizator de active pentru hărți web, conceput pentru a transforma seturile de date spațiale între diferite sisteme de referință de coordonate și formate de fișiere. Proiectul se distinge prin simplificarea geometriei care păstrează topologia, ceea ce reduce numărul de noduri (vertex) menținând în același timp limitele partajate pentru a preveni golurile și suprapunerile. Optimizează în continuare activele pentru web prin cuantificarea coordonatelor și filtrarea atributelor pentru a reduce dimensiunile fișierelor. Sistemul acoperă o gamă largă de capabilități, inclusiv reproiectarea coordonatelor folosind șiruri PROJ și coduri EPSG, și conversia datelor între formate precum Shapefile, GeoJSON, TopoJSON, GeoPackage și KML. Oferă instrumente extinse de procesare a geometriei pentru buffering, clipping, dizolvare și repararea topologiilor, precum și utilitare de gestionare a datelor pentru unirea atributelor, filtrare și transformare. În plus, include funcții de vizualizare pentru generarea de exporturi SVG stilizate, graticule și hărți cu simboluri proporționale. Capabilitățile de procesare spațială pot fi integrate direct în aplicațiile JavaScript și în pipeline-urile de build prin biblioteca sa Node.js.

    Deletes features that share the same identifier as a previous feature to clean datasets.

    JavaScript
    Vezi pe GitHub↗4,133
  • huawei-noah/pretrained-language-modelAvatar huawei-noah

    huawei-noah/Pretrained-Language-Model

    3,163Vezi pe GitHub↗

    Pretrained-Language-Model is a machine learning library and natural language processing toolkit designed for pretraining, tokenizing, and compressing large language models using transformer architectures and specialized optimization techniques. It supports Chinese and multilingual natural language processing tasks, including text classification and conversational response generation. The framework provides specialized capabilities for training large-scale autoregressive and contextual language models, alongside model compression techniques like knowledge distillation and quantization to reduc

    Splits raw text streams into subword tokens using byte-level vocabularies for downstream NLP processing.

    Pythonknowledge-distillationlarge-scale-distributedmodel-compression
    Vezi pe GitHub↗3,163
  1. Home
  2. Artificial Intelligence & ML
  3. Natural Language Processing
  4. Tokenizers
  5. Byte-Level Tokenizers

Explorează sub-etichetele

  • Redundancy Removers2 sub-tag-uriReduces token count by removing redundant tokens identified by a small language model, cutting API costs and latency while preserving task performance. **Distinct from Byte-Level Tokenizers:** Distinct from Byte-Level Tokenizers: focuses on removing redundant tokens for compression, not tokenization methods.
  • TrainingLearns merge rules by iteratively pairing the most frequent adjacent byte sequences in a corpus until a target vocabulary size is reached. **Distinct from Byte-Level Tokenizers:** Distinct from Byte-Level Tokenizers: focuses on the training process to learn merge rules, not the application of an existing tokenizer.