awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to terrifyzhao/bert-utils

Open-source alternatives to Bert Utils

30 open-source projects similar to terrifyzhao/bert-utils, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Bert Utils alternative.

  • codertimo/bert-pytorchالصورة الرمزية لـ codertimo

    codertimo/BERT-pytorch

    6,518عرض على GitHub↗
    Pythonbertlanguage-modelnlp
    عرض على GitHub↗6,518
  • nlpscott/bert-chinese-classification-taskالصورة الرمزية لـ NLPScott

    NLPScott/bert-Chinese-classification-task

    737عرض على GitHub↗

    bert中文分类实践

    Python
    عرض على GitHub↗737
  • socialbird-ailab/bert-classification-tutorialالصورة الرمزية لـ Socialbird-AILab

    Socialbird-AILab/BERT-Classification-Tutorial

    535عرض على GitHub↗

    标注数据,可以说是AI模型训练里最艰巨的一项工作了。自然语言处理的数据标注更是需要投入大量人力。相对计算机视觉的图像标注,文本的标注通常没有准确的标准答案,对句子理解也是因人而异,让这项工作更是难上加难。 但是!谷歌最近发布的BERT大大的解决了这个问题!根据我们的实验,BERT在文本多分类的任务中,能在极小的数据下,带来显著的分类准确率提升。并且,实验主要对比的是仅仅5个月前发布的State of the art 语言模型迁移学习模型 - ULMFiT (https://arxiv.org/abs/1801.06146), 结果有着明显的提升。

    Python
    عرض على GitHub↗535
  • openbmb/bmlistالصورة الرمزية لـ OpenBMB

    OpenBMB/BMList

    345عرض على GitHub↗

    A List of Big Models

    Pythonaiapicode
    عرض على GitHub↗345
  • asyml/texarالصورة الرمزية لـ asyml

    asyml/texar

    2,392عرض على GitHub↗

    Toolkit for Machine Learning, Natural Language Processing, and Text Generation, in TensorFlow. This is part of the CASL project: http://casl-project.ai/

    Pythonbertcasl-projectdata-processing
    عرض على GitHub↗2,392

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • huggingface/transformersالصورة الرمزية لـ huggingface

    huggingface/transformers

    161,630عرض على GitHub↗

    Transformers is a comprehensive library for machine learning that provides a unified interface for training, fine-tuning, and deploying transformer-based models. It supports a wide range of tasks, including text classification, language modeling, question answering, and sequence-to-sequence translation, while offering specialized architectures for both text and vision processing. The framework includes tools for managing the entire model lifecycle, from data preprocessing and tokenization to distributed training and inference. The library features extensive support for model optimization and

    Pythonaudiodeep-learningdeepseek
    عرض على GitHub↗161,630
  • facebookresearch/xlmالصورة الرمزية لـ facebookresearch

    facebookresearch/XLM

    2,930عرض على GitHub↗

    PyTorch original implementation of Cross-lingual Language Model Pretraining.

    Python
    عرض على GitHub↗2,930
  • huggingface/courseالصورة الرمزية لـ huggingface

    huggingface/course

    3,715عرض على GitHub↗

    This project is an educational course and learning curriculum for implementing and fine-tuning transformer models using the Hugging Face ecosystem. It serves as a structured guide and technical walkthrough for processing multimodal data, adapting pre-trained neural networks, and deploying models. The material includes a guide for managing, versioning, and distributing model weights and datasets through a centralized asset hub. It also provides a practical tutorial on adapting models to specific datasets using parameter-efficient methods and an implementation guide for solving natural language

    MDXdeep-learninghacktoberfestnlp
    عرض على GitHub↗3,715
  • openai/gpt-2الصورة الرمزية لـ openai

    openai/gpt-2

    24,967عرض على GitHub↗

    This project is a transformer-based language model and autoregressive text generator designed to predict the next token in a sequence to produce human-like prose and synthetic text. It functions as a large language model that utilizes a transformer architecture to learn linguistic patterns from large datasets for unsupervised multitask learning. The repository provides a distribution of pre-trained weights, enabling natural language processing tasks without requiring additional training. This allows the model to perform zero-shot task generalization by applying learned patterns to new tasks.

    Python
    عرض على GitHub↗24,967
  • morvanzhou/tutorialsالصورة الرمزية لـ MorvanZhou

    MorvanZhou/tutorials

    12,952عرض على GitHub↗

    This repository is a comprehensive collection of instructional guides and practical examples for Python development, focusing on machine learning, data science, and web scraping. It provides implementations for neural networks, reinforcement learning algorithms, and deep learning architectures using PyTorch, alongside detailed manuals for scientific computing and data visualization. The project distinguishes itself by offering specialized tutorials on concurrent programming to optimize CPU performance and guides for setting up Linux development environments. It covers the implementation of ad

    Pythonmachine-learningmultiprocessingneural-network
    عرض على GitHub↗12,952
  • abosamoor/polyglotالصورة الرمزية لـ aboSamoor

    aboSamoor/polyglot

    2,367عرض على GitHub↗

    Multilingual text (NLP) processing toolkit

    Python
    عرض على GitHub↗2,367
  • alexrozanski/llamachatالصورة الرمزية لـ alexrozanski

    alexrozanski/LlamaChat

    1,510عرض على GitHub↗

    Chat with your favourite LLaMA models in a native macOS app

    Swiftaillamallamacpp
    عرض على GitHub↗1,510
  • alexsergivan/transliteratorA

    alexsergivan/transliterator

    0عرض على GitHub↗
    عرض على GitHub↗0
  • alibaba-edu/simple-effective-text-matching-pytorchA

    alibaba-edu/simple-effective-text-matching-pytorch

    0عرض على GitHub↗
    عرض على GitHub↗0
  • aigc-audio/audiogptالصورة الرمزية لـ AIGC-Audio

    AIGC-Audio/AudioGPT

    10,174عرض على GitHub↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Pythonaudiogptmusic
    عرض على GitHub↗10,174
  • abitdodgy/gibranالصورة الرمزية لـ abitdodgy

    abitdodgy/gibran

    65عرض على GitHub↗

    Gibran is an Elixir natural language processor, and a port of WordsCounted.

    Elixir
    عرض على GitHub↗65
  • 7compass/sentimentalالصورة الرمزية لـ 7compass

    7compass/sentimental

    465عرض على GitHub↗

    Simple sentiment analysis with Ruby

    Ruby
    عرض على GitHub↗465
  • ai-shifu/chatallالصورة الرمزية لـ ai-shifu

    ai-shifu/ChatALL

    16,283عرض على GitHub↗

    ChatALL is a desktop application that functions as a multi-model chat client and aggregator for artificial intelligence services. It enables users to send a single prompt to multiple AI models simultaneously, allowing for the side-by-side comparison of generated responses within a unified interface. The application distinguishes itself through a local-first approach to data management, ensuring that all conversation logs and user configurations are stored directly on the user's device. This architecture supports privacy and offline access while providing a centralized system for managing and

    JavaScriptbingchatchatbotchatgpt
    عرض على GitHub↗16,283
  • allenai/mmc4الصورة الرمزية لـ allenai

    allenai/mmc4

    953عرض على GitHub↗

    MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.

    Python
    عرض على GitHub↗953
  • allenai/allennlpالصورة الرمزية لـ allenai

    allenai/allennlp

    11,889عرض على GitHub↗

    AllenNLP is a PyTorch-based research library and deep learning language toolkit designed for developing and training neural network architectures for linguistic tasks. It provides a distributed training system that coordinates data and gradients across multiple GPUs and a framework for integrating pretrained transformer architectures. The system distinguishes itself with a dedicated algorithmic bias mitigation tool used to identify and reduce bias in linguistic model predictions. It also includes model influence analysis to interpret predictions by calculating the influence of specific traini

    Python
    عرض على GitHub↗11,889
  • allenai/scispacyالصورة الرمزية لـ allenai

    allenai/SciSpaCy

    1,968عرض على GitHub↗

    This repository contains custom pipes and models related to using spaCy for scientific documents.

    Python
    عرض على GitHub↗1,968
  • alvations/annotate-questionnaireالصورة الرمزية لـ alvations

    alvations/annotate-questionnaire

    59عرض على GitHub↗

    Summary of Responses to Questionnaire on Annotation Platform https://forms.gle/iZk8kehkjAWmB8xe9

    عرض على GitHub↗59
  • anujvyas/natural-language-processing-projectsالصورة الرمزية لـ anujvyas

    anujvyas/Natural-Language-Processing-Projects

    254عرض على GitHub↗

    This repository consists of all my NLP Projects

    Jupyter Notebook
    عرض على GitHub↗254
  • arc53/docsgptالصورة الرمزية لـ arc53

    arc53/DocsGPT

    17,939عرض على GitHub↗

    DocsGPT is a retrieval-augmented generation platform and private knowledge base used to build AI agents that perform grounded search and analysis. It functions as a multi-model AI orchestrator and enterprise agent builder, allowing for the integration of various local and cloud language models to customize reasoning and text generation. The project provides a visual environment for developing automated assistants using conditional logic and third-party API connectivity. It enables the creation of private AI agents capable of performing enterprise search and detailed document analysis using pr

    Pythonagent-builderagentsai
    عرض على GitHub↗17,939
  • argilla-io/argillaالصورة الرمزية لـ argilla-io

    argilla-io/argilla

    5,015عرض على GitHub↗

    Argilla is a collaborative AI feedback tool and data curation management system. It serves as a human-in-the-loop dataset platform designed to coordinate workforce annotators and domain experts in labeling, rating, and refining data samples for machine learning projects. The platform focuses on large language model dataset curation and reinforcement learning from human feedback workflows. It provides a shared workspace for integrating human expertise into AI development to validate model outputs and correct data errors. The system manages the end-to-end machine learning data pipeline, includ

    Python
    عرض على GitHub↗5,015
  • arongdari/python-topic-modelالصورة الرمزية لـ arongdari

    arongdari/python-topic-model

    374عرض على GitHub↗

    Implementation of various topic models

    Jupyter Notebook
    عرض على GitHub↗374
  • arongdari/topic-model-lecture-noteالصورة الرمزية لـ arongdari

    arongdari/topic-model-lecture-note

    22عرض على GitHub↗

    lecture notes for probabilistic topic models using ipython notebook

    عرض على GitHub↗22
  • artidoro/qloraالصورة الرمزية لـ artidoro

    artidoro/qlora

    10,929عرض على GitHub↗

    This project is a quantized fine-tuning framework for large language models. It implements a low-rank adaptation library and a four-bit quantizer to reduce the GPU memory requirements needed to train large models. The framework utilizes four-bit quantization and low-rank adapters to enable model training on consumer-grade hardware. It further reduces the memory footprint through double quantization and a paged optimizer that offloads states to system RAM. The system supports distributed training across multiple GPUs to handle larger parameter scales and includes utilities for custom dataset

    Jupyter Notebook
    عرض على GitHub↗10,929
  • artificiai/multilingual-latent-dirichlet-allocation-ldaالصورة الرمزية لـ ArtificiAI

    ArtificiAI/Multilingual-Latent-Dirichlet-Allocation-LDA

    83عرض على GitHub↗

    A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

    Pythonclusteringenglishfrench
    عرض على GitHub↗83
  • ahmedbesbes/character-based-cnnA

    ahmedbesbes/character-based-cnn

    0عرض على GitHub↗
    عرض على GitHub↗0