8 مستودعات
Tokenization methods that operate on raw byte sequences to handle diverse vocabularies.
Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Byte-Level Tokenizers. Refine with filters or upvote what's useful.
This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies
Converts raw text into subword units using byte-level sequences to handle diverse languages without requiring language-specific rules.
SentencePiece is a text segmentation engine and tokenization library designed for machine learning workflows. It provides a comprehensive toolkit for transforming raw text into subword units or numerical identifiers, enabling consistent data representation for neural network training and inference. The library supports the training of segmentation models from raw text, allowing for the creation of custom vocabularies tailored to specific domain requirements. The project distinguishes itself through its byte-level encoding and fallback mechanisms, which ensure that every input can be represent
Decomposes unknown characters into UTF-8 byte sequences to ensure full vocabulary coverage without unknown tokens.
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.
Provides a clean implementation of the BPE training algorithm to learn merge rules from text corpora.
LLMLingua is a prompt compression tool that reduces token count in prompts before they are sent to a large language model, cutting API costs and latency while preserving task performance. It operates as an extractive pipeline using a BERT-level Transformer encoder to classify each token for removal based on full bidirectional context from the prompt, retaining only key information and discarding non-essential tokens. The tool is trained through a knowledge distillation process, where a compact compression model learns from an extractive dataset derived from a large language model's output to
Removes redundant tokens identified by a small language model to cut API costs and latency.
Orange3 is a visual data mining platform that provides an interactive canvas for building data analysis workflows without writing code. At its core, it offers a widget-based visual programming environment where users connect configurable components to perform data preprocessing, machine learning model training, statistical evaluation, and interactive visualization. The platform is built on NumPy-backed data tables with domain descriptors that define variable names, types, and roles, and includes a lazy SQL query proxy for working with database tables without loading all data into memory. The
Provides a widget to drop constant attributes and unused categorical values from datasets.
onnxsim هو محسن رسوم بيانية لتعلم الآلة ومبسط للنماذج مصمم لتقليل تعقيد رسوم بيانية حسابات ONNX. يعمل كضاغط للنماذج يستبدل تسلسلات المشغل المعقدة بمخرجات ثابتة مبسطة لتقليل العبء التشغيلي. يحقق المشروع التبسيط من خلال استنتاج الطي الثابت (constant folding)، والذي يستبدل الرسوم البيانية الفرعية للمشغلات الثابتة بموترات ثابتة محسوبة مسبقاً. يستخدم إعادة كتابة الرسوم البيانية القائمة على الأنماط وتحليل الرسوم البيانية للحسابات الثابتة لتحديد وإزالة العقد الزائدة أو العمليات التي لا يمكن الوصول إليها. تغطي الأداة قدرات واسعة لتحسين النماذج، بما في ذلك القضاء على تكرار المشغل وإزالة عقد إعادة التشكيل أو الهوية غير الضرورية. تعمل هذه العمليات على تبسيط تدفق التنفيذ وتقليل بصمة الذاكرة للنموذج.
Eliminates identity operations and unnecessary reshape nodes that do not alter mathematical output.
Mapshaper هي أداة لمعالجة وتبسيط وتحويل البيانات المتجهة الجغرافية، متاحة كواجهة سطر أوامر، وأداة متصفح ويب، ومكتبة Node.js. تعمل كمسقط للإحداثيات، ومحول للبيانات المتجهة، ومحسن لأصول خرائط الويب مصمم لتحويل مجموعات البيانات المكانية بين أنظمة مرجعية إحداثية وتنسيقات ملفات مختلفة. يتميز المشروع بتبسيط الهندسة مع الحفاظ على الطوبولوجيا، مما يقلل من عدد الرؤوس مع الحفاظ على الحدود المشتركة لمنع الفجوات والتداخلات. كما يعمل على تحسين الأصول للويب من خلال تكميم الإحداثيات وتصفية السمات لتقليل أحجام الملفات. يغطي النظام مجموعة واسعة من الإمكانيات، بما في ذلك إعادة إسقاط الإحداثيات باستخدام سلاسل PROJ ورموز EPSG، وتحويل البيانات عبر تنسيقات مثل Shapefile وGeoJSON وTopoJSON وGeoPackage وKML. ويوفر أدوات معالجة هندسية واسعة النطاق للتخزين المؤقت، والقص، والإذابة، وإصلاح الطوبولوجيا، بالإضافة إلى أدوات إدارة البيانات لربط السمات وتصفيتها وتحويلها. بالإضافة إلى ذلك، يتضمن ميزات تصور لتوليد صادرات SVG مصممة، وشبكات إحداثيات، وخرائط رموز متناسبة. يمكن دمج إمكانيات المعالجة المكانية مباشرة في تطبيقات JavaScript وخطوط أنابيب البناء عبر مكتبة Node.js الخاصة به.
Deletes features that share the same identifier as a previous feature to clean datasets.
Pretrained-Language-Model is a machine learning library and natural language processing toolkit designed for pretraining, tokenizing, and compressing large language models using transformer architectures and specialized optimization techniques. It supports Chinese and multilingual natural language processing tasks, including text classification and conversational response generation. The framework provides specialized capabilities for training large-scale autoregressive and contextual language models, alongside model compression techniques like knowledge distillation and quantization to reduc
Splits raw text streams into subword tokens using byte-level vocabularies for downstream NLP processing.