awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to deepset-ai/farm

Open-source alternatives to FARM

30 open-source projects similar to deepset-ai/farm, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best FARM alternative.

  • pytorch/fairseqالصورة الرمزية لـ pytorch

    pytorch/fairseq

    32,228عرض على GitHub↗

    Fairseq is a deep learning research toolkit and sequence-to-sequence framework built on PyTorch. It provides a system for training and deploying models that map input sequences to output sequences, with a primary focus on neural machine translation and speech recognition. The toolkit allows for the generation of text sequences through search algorithms such as beam search and nucleus sampling. It includes capabilities for producing synthetic parallel training data by translating monolingual text using reverse sequence models. The framework supports large scale model training through multi-de

    Python
    عرض على GitHub↗32,228
  • huggingface/transformersالصورة الرمزية لـ huggingface

    huggingface/transformers

    161,630عرض على GitHub↗

    Transformers is a comprehensive library for machine learning that provides a unified interface for training, fine-tuning, and deploying transformer-based models. It supports a wide range of tasks, including text classification, language modeling, question answering, and sequence-to-sequence translation, while offering specialized architectures for both text and vision processing. The framework includes tools for managing the entire model lifecycle, from data preprocessing and tokenization to distributed training and inference. The library features extensive support for model optimization and

    Pythonaudiodeep-learningdeepseek
    عرض على GitHub↗161,630
  • nielsrogge/transformers-tutorialsالصورة الرمزية لـ NielsRogge

    NielsRogge/Transformers-Tutorials

    11,641عرض على GitHub↗

    This is a collection of tutorials and practical demonstrations for implementing machine learning tasks using the HuggingFace Transformers library. It serves as a guide for applying transformer architectures across computer vision, natural language processing, and audio analysis. The repository provides implementation examples for multimodal model deployment, including the combination of text, image, and audio inputs. It includes resources for optimizing pre-trained models through fine-tuning on custom datasets and provides examples for preparing PyTorch datasets by converting raw files into t

    Jupyter Notebookbertgpt-2layoutlm
    عرض على GitHub↗11,641

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • stanfordnlp/stanzaالصورة الرمزية لـ stanfordnlp

    stanfordnlp/stanza

    7,809عرض على GitHub↗

    Stanza is a Python natural language processing library designed for tokenization, lemmatization, and dependency parsing across many human languages using neural models. It provides a neural processing pipeline that converts raw text into structured linguistic data objects, alongside a specialized analyzer for extracting medical insights from clinical and biomedical language. The project includes a wrapper that connects Python scripts to Java-based natural language processing tools and remote annotation servers. This enables a bridge for extracting linguistic annotations and analysis data from

    Pythonartificial-intelligencecorenlpdeep-learning
    عرض على GitHub↗7,809
  • yandexdataschool/nlp_courseالصورة الرمزية لـ yandexdataschool

    yandexdataschool/nlp_course

    10,591عرض على GitHub↗

    YSDA course in Natural Language Processing

    Jupyter Notebook
    عرض على GitHub↗10,591
  • bigartm/bigartmالصورة الرمزية لـ bigartm

    bigartm/bigartm

    674عرض على GitHub↗

    Fast topic modeling platform

    C++
    عرض على GitHub↗674
  • bojone/bert4kerasالصورة الرمزية لـ bojone

    bojone/bert4keras

    5,419عرض على GitHub↗

    bert4keras is a lightweight reimplementation of the BERT transformer architecture for the Keras deep learning framework. It serves as a natural language processing toolkit and transformer model library used for text classification, sequence labeling, and semantic embedding extraction. The framework includes a sequence-to-sequence model system for question answering and text generation, as well as a model inference server to deploy trained transformers as web APIs for real-time predictions. Capabilities cover a broad range of natural language understanding tasks, including reading comprehensi

    Python
    عرض على GitHub↗5,419
  • brikerman/kashgariالصورة الرمزية لـ BrikerMan

    BrikerMan/Kashgari

    2,383عرض على GitHub↗

    Kashgari is a production-level NLP Transfer learning framework built on top of tf.keras for text-labeling and text-classification, includes Word2Vec, BERT, and GPT2 Language Embedding.

    Python
    عرض على GitHub↗2,383
  • chakki-works/chazutsuالصورة الرمزية لـ chakki-works

    chakki-works/chazutsu

    240عرض على GitHub↗

    photo from Kaikado, traditional Japanese chazutsu maker

    Python
    عرض على GitHub↗240
  • chartbeat-labs/textacyالصورة الرمزية لـ chartbeat-labs

    chartbeat-labs/textacy

    2,242عرض على GitHub↗

    NLP, before and after spaCy

    Python
    عرض على GitHub↗2,242
  • codertimo/bert-pytorchالصورة الرمزية لـ codertimo

    codertimo/BERT-pytorch

    6,518عرض على GitHub↗
    Pythonbertlanguage-modelnlp
    عرض على GitHub↗6,518
  • columbia-applied-data-science/rosettaالصورة الرمزية لـ columbia-applied-data-science

    columbia-applied-data-science/rosetta

    207عرض على GitHub↗

    Tools, wrappers, etc... for data science with a concentration on text processing

    Jupyter Notebook
    عرض على GitHub↗207
  • cyberzhg/keras-bertالصورة الرمزية لـ CyberZHG

    CyberZHG/keras-bert

    2,419عرض على GitHub↗

    \中文|English\

    Python
    عرض على GitHub↗2,419
  • datquocnguyen/jptdpالصورة الرمزية لـ datquocnguyen

    datquocnguyen/jPTDP

    156عرض على GitHub↗

    Implementations of joint models for POS tagging and dependency parsing, as described in my papers:

    Python
    عرض على GitHub↗156
  • dbiir/uer-pyالصورة الرمزية لـ dbiir

    dbiir/UER-py

    3,108عرض على GitHub↗

    Open Source Pre-training Model Framework in PyTorch & Pre-trained Model Zoo

    Python
    عرض على GitHub↗3,108
  • deepset-ai/haystackالصورة الرمزية لـ deepset-ai

    deepset-ai/haystack

    24,253عرض على GitHub↗

    Haystack is an orchestration framework designed for building complex search and generative AI pipelines. It functions as an agentic workflow engine, enabling the construction of automated sequences that allow AI agents to perform multi-step reasoning and data analysis. The framework utilizes a modular, component-based architecture that connects processing steps into directed acyclic graphs. By employing a provider-agnostic integration layer, it decouples core logic from specific external AI services and vector databases, allowing for the flexible exchange of underlying technologies. This desi

    MDXagentagentsai
    عرض على GitHub↗24,253
  • dhlee347/pytorchic-bertالصورة الرمزية لـ dhlee347

    dhlee347/pytorchic-bert

    599عرض على GitHub↗

    Pytorch Implementation of Google BERT

    Python
    عرض على GitHub↗599
  • dmlc/gluon-nlpالصورة الرمزية لـ dmlc

    dmlc/gluon-nlp

    2,546عرض على GitHub↗

    .. raw:: html

    Python
    عرض على GitHub↗2,546
  • dreamgonfly/bert-pytorchالصورة الرمزية لـ dreamgonfly

    dreamgonfly/BERT-pytorch

    110عرض على GitHub↗

    PyTorch implementation of BERT in "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" (https://arxiv.org/abs/1810.04805)

    Python
    عرض على GitHub↗110
  • explosion/spacyالصورة الرمزية لـ explosion

    explosion/spaCy

    33,688عرض على GitHub↗

    spaCy is a Python natural language processing framework designed for industrial-scale text processing. It converts raw text into structured data for machine learning pipelines through a combination of statistical language model trainers, transformer-based text processors, and syntactic dependency parsers. The project enables the integration of pretrained transformer architectures to perform complex linguistic analysis and multi-task learning. It also provides a specialized system for neural named entity recognition to identify and categorize key entities within text. The framework covers a b

    Pythonaiartificial-intelligencecython
    عرض على GitHub↗33,688
  • explosion/spacy-transformersالصورة الرمزية لـ explosion

    explosion/spacy-transformers

    1,406عرض على GitHub↗

    This package provides spaCy components and architectures to use transformer models via Hugging Face's transformers in spaCy. The result is convenient access to state-of-the-art transformer architectures, such as BERT, GPT-2, XLNet, etc.

    Python
    عرض على GitHub↗1,406
  • facebookresearch/mmbtالصورة الرمزية لـ facebookresearch

    facebookresearch/mmbt

    257عرض على GitHub↗

    MMBT is the accompanying code repository for the paper titled, "Supervised Multimodal Bitransformers for Classifying Images and Text" by Douwe Kiela, Suvrat Bhooshan, Hamed Firooz, Ethan Perez and Davide Testuggine.

    Python
    عرض على GitHub↗257
  • ggerganov/llama.cppالصورة الرمزية لـ ggerganov

    ggerganov/llama.cpp

    116,912عرض على GitHub↗

    llama.cpp is a high-performance C++ inference engine and runtime for executing large language models locally across various hardware architectures. It provides the core components for local model execution, including a dedicated model quantizer for compressing weights into the GGUF format and a system for generating text embeddings for semantic search. The project distinguishes itself through specialized memory and execution optimizations, such as block-wise weight quantization to reduce memory footprints and memory-mapped model loading. It supports structured text generation by using formal

    C++
    عرض على GitHub↗116,912
  • google-research/bertالصورة الرمزية لـ google-research

    google-research/bert

    39,869عرض على GitHub↗

    This project is a transformer-based language model and natural language processing toolkit designed to generate deep contextual representations of text. By utilizing a transformer-based encoder architecture, the system processes input sequences through stacked self-attention layers to capture the semantic meaning of tokens based on their surrounding sentence structure. The model distinguishes itself through bidirectional contextual processing, which analyzes text in both directions simultaneously, and masked language modeling, which trains the system by predicting hidden tokens within a seque

    Pythongooglenatural-language-processingnatural-language-understanding
    عرض على GitHub↗39,869
  • guotong1988/bert-tensorflowG

    guotong1988/BERT-tensorflow

    0عرض على GitHub↗
    عرض على GitHub↗0
  • huggingface/tokenizersالصورة الرمزية لـ huggingface

    huggingface/tokenizers

    10,825عرض على GitHub↗

    This project is a high-performance library for converting raw text into tokens and IDs for machine learning models. It functions as a fast text encoder and a text preprocessing pipeline designed to transform strings into numerical representations with high throughput for research and production. The library includes a subword tokenizer trainer used to analyze text datasets and create custom vocabularies using algorithms such as byte-pair encoding and wordpiece. It provides capabilities for subword vocabulary training and text alignment, allowing character offsets to be tracked during normaliz

    Rustbertgptlanguage-model
    عرض على GitHub↗10,825
  • innodatalabs/tbertالصورة الرمزية لـ innodatalabs

    innodatalabs/tbert

    17عرض على GitHub↗

    BERT model converted to PyTorch.

    Python
    عرض على GitHub↗17
  • jasonkessler/scattertextالصورة الرمزية لـ JasonKessler

    JasonKessler/scattertext

    2,330عرض على GitHub↗

    Beautiful visualizations of how language differs among document types.

    Pythoncomputational-social-scienced3eda
    عرض على GitHub↗2,330
  • kaushaltrivedi/fast-bertالصورة الرمزية لـ kaushaltrivedi

    kaushaltrivedi/fast-bert

    1,916عرض على GitHub↗

    New - Learning Rate Finder for Text Classification Training (borrowed with thanks from https://github.com/davidtvs/pytorch-lr-finder)

    Python
    عرض على GitHub↗1,916
  • kimiyoung/transformer-xlالصورة الرمزية لـ kimiyoung

    kimiyoung/transformer-xl

    3,703عرض على GitHub↗

    This project is an implementation of the Transformer-XL language model, a neural network architecture designed for long-context language modeling. It provides frameworks for training and deploying models that capture long-term dependencies and relationships in text sequences that extend beyond a fixed context window. The implementation supports both PyTorch and TensorFlow, allowing for distributed training across multiple GPUs and host nodes. It employs a recurrent mechanism to maintain coherence in extended sequences, utilizing segment-level recurrence and state-based memory reuse. The code

    Python
    عرض على GitHub↗3,703