awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to kangfend/bahasa

Open-source alternatives to Bahasa

30 open-source projects similar to kangfend/bahasa, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Bahasa alternative.

  • 01walid/goarabic01walid 的头像

    01walid/goarabic

    117在 GitHub 上查看↗

    A Go Lang package for dealing with Arabic text.

    Go
    在 GitHub 上查看↗117
  • adbar/german-nlpadbar 的头像

    adbar/German-NLP

    526在 GitHub 上查看↗

    Curated list of open-access/open-source/off-the-shelf resources and tools developed with a particular focus on German

    在 GitHub 上查看↗526
  • adobe/nlp-cubeadobe 的头像

    adobe/NLP-Cube

    562在 GitHub 上查看↗

    05 August 2021 - We are releasing version 3.0 of NLPCube and models and introducing FLAVOURS. This is a major update, but we did our best to maintain the same API, so previous implementation will not crash. The supported language list is smaller, but you can open an issue for unsupported…

    HTML
    在 GitHub 上查看↗562
  • alexandrainst/danlpalexandrainst 的头像

    alexandrainst/danlp

    209在 GitHub 上查看↗

    Part of Speech Tagging | Dependency Parsing Named Entity Recognition | Named Entity Disambiguation | Coreference Resolution Sentiment Analysis | Hatespeech Detection Embeddings | Datasets | Tutorials

    Python
    在 GitHub 上查看↗209
  • alirezatheh/perkeAlirezaTheH 的头像

    AlirezaTheH/perke

    73在 GitHub 上查看↗

    Perke is a Python keyphrase extraction package for Persian language. It provides an end-to-end keyphrase extraction pipeline in which each component can be easily modified or extended to develop new models.

    Python
    在 GitHub 上查看↗73

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • amir-zeldes/rftokenizeramir-zeldes 的头像

    amir-zeldes/RFTokenizer

    31在 GitHub 上查看↗

    A character-wise tokenizer for morphologically rich languages

    Lex
    在 GitHub 上查看↗31
  • aziz/virastaraziz 的头像

    aziz/virastar

    88在 GitHub 上查看↗

    #ویراستار نوشته‌های فارسی شما را ویرایش می‌کند

    Ruby
    在 GitHub 上查看↗88
  • botcenter/spanishsent2vecBotCenter 的头像

    BotCenter/spanishSent2Vec

    4在 GitHub 上查看↗

    Spanish Sentence Embeddings trained using sent2vec on the Spanish Unannotated Corpora.

    在 GitHub 上查看↗4
  • botcenter/spanishwordembeddingsBotCenter 的头像

    BotCenter/spanishWordEmbeddings

    9在 GitHub 上查看↗

    Spanish words embeddings computed using fastText on the Spanish Unannotated Corpora.

    在 GitHub 上查看↗9
  • calmdownkarm/sivareddydependencyparserCalmDownKarm 的头像

    CalmDownKarm/sivareddydependencyparser

    0在 GitHub 上查看↗

    Your input file should have the extension .input.txt e.g. hindi.input.txt To dependency tag your input file run "make .output" e.g.

    Lex
    在 GitHub 上查看↗0
  • dccuchile/betodccuchile 的头像

    dccuchile/beto

    505在 GitHub 上查看↗

    BETO is a BERT model trained on a big Spanish corpus. BETO is of size similar to a BERT-Base and was trained with the Whole Word Masking technique. Below you find Tensorflow and Pytorch checkpoints for the uncased and cased versions, as well as some results for Spanish benchmarks comparing BETO…

    在 GitHub 上查看↗505
  • dccuchile/spanish-word-embeddingsdccuchile 的头像

    dccuchile/spanish-word-embeddings

    365在 GitHub 上查看↗

    Below you find links to Spanish word embeddings computed with different methods and from different corpora. Whenever it is possible, a description of the parameters used to compute the embeddings is included, together with simple statistics of the vectors, vocabulary, and description of the…

    在 GitHub 上查看↗365
  • ejtaal/jsastemejtaal 的头像

    ejtaal/jsastem

    26在 GitHub 上查看↗

    JSASTEM - JavaScript Arabic Stemmer

    JavaScript
    在 GitHub 上查看↗26
  • fighting41love/funnlpfighting41love 的头像

    fighting41love/funNLP

    81,299在 GitHub 上查看↗

    This project is a community-driven knowledge base and curated repository focused on natural language processing and large language model development. It serves as a centralized index for high-quality tools, libraries, and research materials, organizing technical resources into structured, version-controlled documentation to assist developers in navigating the evolving artificial intelligence ecosystem. The repository distinguishes itself by acting as an aggregator for AI model evaluation and benchmarking. It provides access to tools that enable the simultaneous comparison of multiple conversa

    Python
    在 GitHub 上查看↗81,299
  • fnielsen/awesome-danishfnielsen 的头像

    fnielsen/awesome-danish

    195在 GitHub 上查看↗

    A curated list of awesome resources for Danish language technology

    在 GitHub 上查看↗195
  • fudannlp/fnlpFudanNLP 的头像

    FudanNLP/fnlp

    2,690在 GitHub 上查看↗

    FudanNLP (FNLP)

    Java
    在 GitHub 上查看↗2,690
  • fxsjy/jiebafxsjy 的头像

    fxsjy/jieba

    35,027在 GitHub 上查看↗

    This project is a Chinese text segmentation library and tokenizer designed to split Chinese sentences into individual words. It serves as a natural language processing tool for splitting characters into words, tagging parts of speech, and extracting keywords using statistical analysis. The library distinguishes itself through support for custom dictionary configuration and vocabulary file management, allowing users to override default segmentation rules for domain-specific accuracy. It also includes a TF-IDF keyword extractor to identify significant words and core topics within documents. Th

    Python
    在 GitHub 上查看↗35,027
  • galuhsahid/indonesian-word-embeddinggaluhsahid 的头像

    galuhsahid/indonesian-word-embedding

    20在 GitHub 上查看↗

    A web application that demonstrates Indonesian word embedding, inspired by Word embedding demo.

    JavaScript
    在 GitHub 上查看↗20
  • goru001/inltkgoru001 的头像

    goru001/inltk

    840在 GitHub 上查看↗

    iNLTK aims to provide out of the box support for various NLP tasks that an application developer might need for Indic languages.

    Python
    在 GitHub 上查看↗840
  • hankcs/hanlphankcs 的头像

    hankcs/HanLP

    36,413在 GitHub 上查看↗

    HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion. The project distinguishes itself through a dedicated focus on Chinese linguistic structures, including a specialized script converter for transforming text between Simplified Chinese, Traditional Chinese, and Pinyin. It further supports domain-specific model training to improve the recognition of professional t

    Pythondependency-parserhanlpnamed-entity-recognition
    在 GitHub 上查看↗36,413
  • ictrc/parsivarICTRC 的头像

    ICTRC/Parsivar

    247在 GitHub 上查看↗

    parsivar

    Python
    在 GitHub 上查看↗247
  • isnowfy/snownlpisnowfy 的头像

    isnowfy/snownlp

    6,631在 GitHub 上查看↗

    SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation, sentiment analysis, document classification, and phonetic transliteration. The library includes capabilities for training and saving custom machine learning models for tokenization and sentiment analysis using raw training datasets. It covers a range of linguistic processing areas, including parts of speech tagging, sentence splitting, and text similarity measurement. The toolkit also provides utilities for extracting key information through text summarization and calculating word im

    Python
    在 GitHub 上查看↗6,631
  • jfreddypuentes/spanlpjfreddypuentes 的头像

    jfreddypuentes/spanlp

    41在 GitHub 上查看↗

    spanlp es una librería escrita en Python para detectar, censurar y limpiar groserías, vulgaridades, palabras de odio, racismo, xenofobia y bullying en textos escritos en Español .

    Python
    在 GitHub 上查看↗41
  • jonsafari/perstemjonsafari 的头像

    jonsafari/perstem

    19在 GitHub 上查看↗

    Persian (Farsi) stemmer, morphological analyzer, transliterator, and partial part-of-speech tagger. Input may be encoded as Perso-Arabic script UTF-8, ISIRI 3342, Windows-1256, SGML/HTML/XML-style numeric character references (ncr), or dehdari-transliterated latin-script text. Use the -i flag to…

    Perl
    在 GitHub 上查看↗19
  • kenjiroai/synthaiKenjiroAI 的头像

    KenjiroAI/SynThai

    41在 GitHub 上查看↗

    Thai Word Segmentation and Part-of-Speech Tagging with Deep Learning

    Python
    在 GitHub 上查看↗41
  • ksopyla/awesome-nlp-polishksopyla 的头像

    ksopyla/awesome-nlp-polish

    308在 GitHub 上查看↗

    A curated list of resources dedicated to Natural Language Processing (NLP) in polish. Models, tools, datasets.

    在 GitHub 上查看↗308
  • mikahama/uralicnlpmikahama 的头像

    mikahama/uralicNLP

    98在 GitHub 上查看↗

    Natural language processing for many languages

    Python
    在 GitHub 上查看↗98
  • narimann2/parsianalyzerNarimanN2 的头像

    NarimanN2/ParsiAnalyzer

    166在 GitHub 上查看↗

    Persian Analyzer for Elasticsearch.

    Java
    在 GitHub 上查看↗166
  • phuonglh/vn.vitkphuonglh 的头像

    phuonglh/vn.vitk

    218在 GitHub 上查看↗

    NOTE: This repos is now obsolete. Interested programmers should consider to use the new repo vlp (github.com/phuonglh/vlp) We have preferred using Scala instead of Java since 2016.

    Java
    在 GitHub 上查看↗218
  • proycon/python-frogproycon 的头像

    proycon/python-frog

    49在 GitHub 上查看↗

    Python bindings to the dutch NLP tool Frog (pos tagger, lemmatiser, NER tagger, morphological analysis, shallow parser, dependency parser)

    Cython
    在 GitHub 上查看↗49