awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to urduhack/urduhack

Open-source alternatives to Urduhack

30 open-source projects similar to urduhack/urduhack, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Urduhack alternative.

  • 01walid/goarabicAvatar de 01walid

    01walid/goarabic

    117Ver en GitHub↗

    A Go Lang package for dealing with Arabic text.

    Go
    Ver en GitHub↗117
  • adbar/german-nlpAvatar de adbar

    adbar/German-NLP

    526Ver en GitHub↗

    Curated list of open-access/open-source/off-the-shelf resources and tools developed with a particular focus on German

    Ver en GitHub↗526
  • adobe/nlp-cubeAvatar de adobe

    adobe/NLP-Cube

    562Ver en GitHub↗

    05 August 2021 - We are releasing version 3.0 of NLPCube and models and introducing FLAVOURS. This is a major update, but we did our best to maintain the same API, so previous implementation will not crash. The supported language list is smaller, but you can open an issue for unsupported…

    HTML
    Ver en GitHub↗562
  • alexandrainst/danlpAvatar de alexandrainst

    alexandrainst/danlp

    209Ver en GitHub↗

    Part of Speech Tagging | Dependency Parsing Named Entity Recognition | Named Entity Disambiguation | Coreference Resolution Sentiment Analysis | Hatespeech Detection Embeddings | Datasets | Tutorials

    Python
    Ver en GitHub↗209
  • alirezatheh/perkeAvatar de AlirezaTheH

    AlirezaTheH/perke

    73Ver en GitHub↗

    Perke is a Python keyphrase extraction package for Persian language. It provides an end-to-end keyphrase extraction pipeline in which each component can be easily modified or extended to develop new models.

    Python
    Ver en GitHub↗73

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Find more with AI search
  • amir-zeldes/rftokenizerAvatar de amir-zeldes

    amir-zeldes/RFTokenizer

    31Ver en GitHub↗

    A character-wise tokenizer for morphologically rich languages

    Lex
    Ver en GitHub↗31
  • aziz/virastarAvatar de aziz

    aziz/virastar

    88Ver en GitHub↗

    #ویراستار نوشته‌های فارسی شما را ویرایش می‌کند

    Ruby
    Ver en GitHub↗88
  • botcenter/spanishsent2vecAvatar de BotCenter

    BotCenter/spanishSent2Vec

    4Ver en GitHub↗

    Spanish Sentence Embeddings trained using sent2vec on the Spanish Unannotated Corpora.

    Ver en GitHub↗4
  • botcenter/spanishwordembeddingsAvatar de BotCenter

    BotCenter/spanishWordEmbeddings

    9Ver en GitHub↗

    Spanish words embeddings computed using fastText on the Spanish Unannotated Corpora.

    Ver en GitHub↗9
  • calmdownkarm/sivareddydependencyparserAvatar de CalmDownKarm

    CalmDownKarm/sivareddydependencyparser

    0Ver en GitHub↗

    Your input file should have the extension .input.txt e.g. hindi.input.txt To dependency tag your input file run "make .output" e.g.

    Lex
    Ver en GitHub↗0
  • dccuchile/betoAvatar de dccuchile

    dccuchile/beto

    505Ver en GitHub↗

    BETO is a BERT model trained on a big Spanish corpus. BETO is of size similar to a BERT-Base and was trained with the Whole Word Masking technique. Below you find Tensorflow and Pytorch checkpoints for the uncased and cased versions, as well as some results for Spanish benchmarks comparing BETO…

    Ver en GitHub↗505
  • dccuchile/spanish-word-embeddingsAvatar de dccuchile

    dccuchile/spanish-word-embeddings

    365Ver en GitHub↗

    Below you find links to Spanish word embeddings computed with different methods and from different corpora. Whenever it is possible, a description of the parameters used to compute the embeddings is included, together with simple statistics of the vectors, vocabulary, and description of the…

    Ver en GitHub↗365
  • ejtaal/jsastemAvatar de ejtaal

    ejtaal/jsastem

    26Ver en GitHub↗

    JSASTEM - JavaScript Arabic Stemmer

    JavaScript
    Ver en GitHub↗26
  • fighting41love/funnlpAvatar de fighting41love

    fighting41love/funNLP

    81,299Ver en GitHub↗

    This project is a community-driven knowledge base and curated repository focused on natural language processing and large language model development. It serves as a centralized index for high-quality tools, libraries, and research materials, organizing technical resources into structured, version-controlled documentation to assist developers in navigating the evolving artificial intelligence ecosystem. The repository distinguishes itself by acting as an aggregator for AI model evaluation and benchmarking. It provides access to tools that enable the simultaneous comparison of multiple conversa

    Python
    Ver en GitHub↗81,299
  • fnielsen/awesome-danishAvatar de fnielsen

    fnielsen/awesome-danish

    195Ver en GitHub↗

    A curated list of awesome resources for Danish language technology

    Ver en GitHub↗195
  • fudannlp/fnlpAvatar de FudanNLP

    FudanNLP/fnlp

    2,690Ver en GitHub↗

    FudanNLP (FNLP)

    Java
    Ver en GitHub↗2,690
  • fxsjy/jiebaAvatar de fxsjy

    fxsjy/jieba

    35,027Ver en GitHub↗

    This project is a Chinese text segmentation library and tokenizer designed to split Chinese sentences into individual words. It serves as a natural language processing tool for splitting characters into words, tagging parts of speech, and extracting keywords using statistical analysis. The library distinguishes itself through support for custom dictionary configuration and vocabulary file management, allowing users to override default segmentation rules for domain-specific accuracy. It also includes a TF-IDF keyword extractor to identify significant words and core topics within documents. Th

    Python
    Ver en GitHub↗35,027
  • galuhsahid/indonesian-word-embeddingAvatar de galuhsahid

    galuhsahid/indonesian-word-embedding

    20Ver en GitHub↗

    A web application that demonstrates Indonesian word embedding, inspired by Word embedding demo.

    JavaScript
    Ver en GitHub↗20
  • goru001/inltkAvatar de goru001

    goru001/inltk

    840Ver en GitHub↗

    iNLTK aims to provide out of the box support for various NLP tasks that an application developer might need for Indic languages.

    Python
    Ver en GitHub↗840
  • hankcs/hanlpAvatar de hankcs

    hankcs/HanLP

    36,413Ver en GitHub↗

    HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion. The project distinguishes itself through a dedicated focus on Chinese linguistic structures, including a specialized script converter for transforming text between Simplified Chinese, Traditional Chinese, and Pinyin. It further supports domain-specific model training to improve the recognition of professional t

    Pythondependency-parserhanlpnamed-entity-recognition
    Ver en GitHub↗36,413
  • ictrc/parsivarAvatar de ICTRC

    ICTRC/Parsivar

    247Ver en GitHub↗

    parsivar

    Python
    Ver en GitHub↗247
  • isnowfy/snownlpAvatar de isnowfy

    isnowfy/snownlp

    6,631Ver en GitHub↗

    SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation, sentiment analysis, document classification, and phonetic transliteration. The library includes capabilities for training and saving custom machine learning models for tokenization and sentiment analysis using raw training datasets. It covers a range of linguistic processing areas, including parts of speech tagging, sentence splitting, and text similarity measurement. The toolkit also provides utilities for extracting key information through text summarization and calculating word im

    Python
    Ver en GitHub↗6,631
  • jfreddypuentes/spanlpAvatar de jfreddypuentes

    jfreddypuentes/spanlp

    41Ver en GitHub↗

    spanlp es una librería escrita en Python para detectar, censurar y limpiar groserías, vulgaridades, palabras de odio, racismo, xenofobia y bullying en textos escritos en Español .

    Python
    Ver en GitHub↗41
  • jonsafari/perstemAvatar de jonsafari

    jonsafari/perstem

    19Ver en GitHub↗

    Persian (Farsi) stemmer, morphological analyzer, transliterator, and partial part-of-speech tagger. Input may be encoded as Perso-Arabic script UTF-8, ISIRI 3342, Windows-1256, SGML/HTML/XML-style numeric character references (ncr), or dehdari-transliterated latin-script text. Use the -i flag to…

    Perl
    Ver en GitHub↗19
  • kangfend/bahasaAvatar de kangfend

    kangfend/bahasa

    20Ver en GitHub↗

    BAHASA

    Python
    Ver en GitHub↗20
  • kenjiroai/synthaiAvatar de KenjiroAI

    KenjiroAI/SynThai

    41Ver en GitHub↗

    Thai Word Segmentation and Part-of-Speech Tagging with Deep Learning

    Python
    Ver en GitHub↗41
  • ksopyla/awesome-nlp-polishAvatar de ksopyla

    ksopyla/awesome-nlp-polish

    308Ver en GitHub↗

    A curated list of resources dedicated to Natural Language Processing (NLP) in polish. Models, tools, datasets.

    Ver en GitHub↗308
  • mikahama/uralicnlpAvatar de mikahama

    mikahama/uralicNLP

    98Ver en GitHub↗

    Natural language processing for many languages

    Python
    Ver en GitHub↗98
  • narimann2/parsianalyzerAvatar de NarimanN2

    NarimanN2/ParsiAnalyzer

    166Ver en GitHub↗

    Persian Analyzer for Elasticsearch.

    Java
    Ver en GitHub↗166
  • phuonglh/vn.vitkAvatar de phuonglh

    phuonglh/vn.vitk

    218Ver en GitHub↗

    NOTE: This repos is now obsolete. Interested programmers should consider to use the new repo vlp (github.com/phuonglh/vlp) We have preferred using Scala instead of Java since 2016.

    Java
    Ver en GitHub↗218