awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 مستودعات

Awesome GitHub RepositoriesLLM-Based Classifiers

Assigns labels to text using a configured large language model, without requiring a full extraction result.

Distinct from Text Classifiers: Distinct from Text Classifiers: uses LLMs for classification rather than traditional ML models or rule-based approaches.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · LLM-Based Classifiers. Refine with filters or upvote what's useful.

Awesome LLM-Based Classifiers GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • kreuzberg-dev/kreuzbergالصورة الرمزية لـ kreuzberg-dev

    kreuzberg-dev/kreuzberg

    8,527عرض على GitHub↗

    Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo

    Assigns labels to a single piece of plain text using a configured LLM, without requiring an extraction result.

    Rustdocument-intelligenceelixirffi
    عرض على GitHub↗8,527
  • datajuicer/data-juicerالصورة الرمزية لـ datajuicer

    datajuicer/data-juicer

    6,574عرض على GitHub↗

    Data-Juicer is an open-source framework for cleaning, filtering, deduplicating, and transforming multimodal datasets to prepare them for training large language and vision models. It functions as a distributed data pipeline engine that runs processing jobs across Ray clusters, handling billions of samples with automatic operator fusion and adaptive parallelism. The framework provides a library of operators that leverage large language models for semantic extraction, filtering, and data synthesis within processing pipelines. The project distinguishes itself through a YAML-based data recipe sys

    Runs a quality classifier on web-crawled text to filter low-quality samples.

    Pythondatadata-analysisdata-pipeline
    عرض على GitHub↗6,574
  1. Home
  2. Artificial Intelligence & ML
  3. Text Classifiers
  4. LLM-Based Classifiers