awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to urduhack/urduhack

Open-source alternatives to Urduhack

30 open-source projects similar to urduhack/urduhack, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Urduhack alternative.

  • 01walid/goarabicAvatar 01walid

    01walid/goarabic

    117Vezi pe GitHub↗

    A Go Lang package for dealing with Arabic text.

    Go
    Vezi pe GitHub↗117
  • adbar/german-nlpAvatar adbar

    adbar/German-NLP

    526Vezi pe GitHub↗

    Curated list of open-access/open-source/off-the-shelf resources and tools developed with a particular focus on German

    Vezi pe GitHub↗526
  • adobe/nlp-cubeAvatar adobe

    adobe/NLP-Cube

    562Vezi pe GitHub↗

    05 August 2021 - We are releasing version 3.0 of NLPCube and models and introducing FLAVOURS. This is a major update, but we did our best to maintain the same API, so previous implementation will not crash. The supported language list is smaller, but you can open an issue for unsupported…

    HTML
    Vezi pe GitHub↗562
  • alexandrainst/danlpAvatar alexandrainst

    alexandrainst/danlp

    209Vezi pe GitHub↗

    Part of Speech Tagging | Dependency Parsing Named Entity Recognition | Named Entity Disambiguation | Coreference Resolution Sentiment Analysis | Hatespeech Detection Embeddings | Datasets | Tutorials

    Python
    Vezi pe GitHub↗209
  • alirezatheh/perkeAvatar AlirezaTheH

    AlirezaTheH/perke

    73Vezi pe GitHub↗

    Perke is a Python keyphrase extraction package for Persian language. It provides an end-to-end keyphrase extraction pipeline in which each component can be easily modified or extended to develop new models.

    Python
    Vezi pe GitHub↗73

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Find more with AI search
  • amir-zeldes/rftokenizerAvatar amir-zeldes

    amir-zeldes/RFTokenizer

    31Vezi pe GitHub↗

    A character-wise tokenizer for morphologically rich languages

    Lex
    Vezi pe GitHub↗31
  • aziz/virastarAvatar aziz

    aziz/virastar

    88Vezi pe GitHub↗

    #ویراستار نوشته‌های فارسی شما را ویرایش می‌کند

    Ruby
    Vezi pe GitHub↗88
  • botcenter/spanishsent2vecAvatar BotCenter

    BotCenter/spanishSent2Vec

    4Vezi pe GitHub↗

    Spanish Sentence Embeddings trained using sent2vec on the Spanish Unannotated Corpora.

    Vezi pe GitHub↗4
  • botcenter/spanishwordembeddingsAvatar BotCenter

    BotCenter/spanishWordEmbeddings

    9Vezi pe GitHub↗

    Spanish words embeddings computed using fastText on the Spanish Unannotated Corpora.

    Vezi pe GitHub↗9
  • calmdownkarm/sivareddydependencyparserAvatar CalmDownKarm

    CalmDownKarm/sivareddydependencyparser

    0Vezi pe GitHub↗

    Your input file should have the extension .input.txt e.g. hindi.input.txt To dependency tag your input file run "make .output" e.g.

    Lex
    Vezi pe GitHub↗0
  • dccuchile/betoAvatar dccuchile

    dccuchile/beto

    505Vezi pe GitHub↗

    BETO is a BERT model trained on a big Spanish corpus. BETO is of size similar to a BERT-Base and was trained with the Whole Word Masking technique. Below you find Tensorflow and Pytorch checkpoints for the uncased and cased versions, as well as some results for Spanish benchmarks comparing BETO…

    Vezi pe GitHub↗505
  • dccuchile/spanish-word-embeddingsAvatar dccuchile

    dccuchile/spanish-word-embeddings

    365Vezi pe GitHub↗

    Below you find links to Spanish word embeddings computed with different methods and from different corpora. Whenever it is possible, a description of the parameters used to compute the embeddings is included, together with simple statistics of the vectors, vocabulary, and description of the…

    Vezi pe GitHub↗365
  • ejtaal/jsastemAvatar ejtaal

    ejtaal/jsastem

    26Vezi pe GitHub↗

    JSASTEM - JavaScript Arabic Stemmer

    JavaScript
    Vezi pe GitHub↗26
  • fighting41love/funnlpAvatar fighting41love

    fighting41love/funNLP

    81,299Vezi pe GitHub↗

    This project is a community-driven knowledge base and curated repository focused on natural language processing and large language model development. It serves as a centralized index for high-quality tools, libraries, and research materials, organizing technical resources into structured, version-controlled documentation to assist developers in navigating the evolving artificial intelligence ecosystem. The repository distinguishes itself by acting as an aggregator for AI model evaluation and benchmarking. It provides access to tools that enable the simultaneous comparison of multiple conversa

    Python
    Vezi pe GitHub↗81,299
  • fnielsen/awesome-danishAvatar fnielsen

    fnielsen/awesome-danish

    195Vezi pe GitHub↗

    A curated list of awesome resources for Danish language technology

    Vezi pe GitHub↗195
  • fudannlp/fnlpAvatar FudanNLP

    FudanNLP/fnlp

    2,690Vezi pe GitHub↗

    FudanNLP (FNLP)

    Java
    Vezi pe GitHub↗2,690
  • fxsjy/jiebaAvatar fxsjy

    fxsjy/jieba

    35,027Vezi pe GitHub↗

    This project is a Chinese text segmentation library and tokenizer designed to split Chinese sentences into individual words. It serves as a natural language processing tool for splitting characters into words, tagging parts of speech, and extracting keywords using statistical analysis. The library distinguishes itself through support for custom dictionary configuration and vocabulary file management, allowing users to override default segmentation rules for domain-specific accuracy. It also includes a TF-IDF keyword extractor to identify significant words and core topics within documents. Th

    Python
    Vezi pe GitHub↗35,027
  • galuhsahid/indonesian-word-embeddingAvatar galuhsahid

    galuhsahid/indonesian-word-embedding

    20Vezi pe GitHub↗

    A web application that demonstrates Indonesian word embedding, inspired by Word embedding demo.

    JavaScript
    Vezi pe GitHub↗20
  • goru001/inltkAvatar goru001

    goru001/inltk

    840Vezi pe GitHub↗

    iNLTK aims to provide out of the box support for various NLP tasks that an application developer might need for Indic languages.

    Python
    Vezi pe GitHub↗840
  • hankcs/hanlpAvatar hankcs

    hankcs/HanLP

    36,413Vezi pe GitHub↗

    HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion. The project distinguishes itself through a dedicated focus on Chinese linguistic structures, including a specialized script converter for transforming text between Simplified Chinese, Traditional Chinese, and Pinyin. It further supports domain-specific model training to improve the recognition of professional t

    Pythondependency-parserhanlpnamed-entity-recognition
    Vezi pe GitHub↗36,413
  • ictrc/parsivarAvatar ICTRC

    ICTRC/Parsivar

    247Vezi pe GitHub↗

    parsivar

    Python
    Vezi pe GitHub↗247
  • isnowfy/snownlpAvatar isnowfy

    isnowfy/snownlp

    6,631Vezi pe GitHub↗

    SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation, sentiment analysis, document classification, and phonetic transliteration. The library includes capabilities for training and saving custom machine learning models for tokenization and sentiment analysis using raw training datasets. It covers a range of linguistic processing areas, including parts of speech tagging, sentence splitting, and text similarity measurement. The toolkit also provides utilities for extracting key information through text summarization and calculating word im

    Python
    Vezi pe GitHub↗6,631
  • jfreddypuentes/spanlpAvatar jfreddypuentes

    jfreddypuentes/spanlp

    41Vezi pe GitHub↗

    spanlp es una librería escrita en Python para detectar, censurar y limpiar groserías, vulgaridades, palabras de odio, racismo, xenofobia y bullying en textos escritos en Español .

    Python
    Vezi pe GitHub↗41
  • jonsafari/perstemAvatar jonsafari

    jonsafari/perstem

    19Vezi pe GitHub↗

    Persian (Farsi) stemmer, morphological analyzer, transliterator, and partial part-of-speech tagger. Input may be encoded as Perso-Arabic script UTF-8, ISIRI 3342, Windows-1256, SGML/HTML/XML-style numeric character references (ncr), or dehdari-transliterated latin-script text. Use the -i flag to…

    Perl
    Vezi pe GitHub↗19
  • kangfend/bahasaAvatar kangfend

    kangfend/bahasa

    20Vezi pe GitHub↗

    BAHASA

    Python
    Vezi pe GitHub↗20
  • kenjiroai/synthaiAvatar KenjiroAI

    KenjiroAI/SynThai

    41Vezi pe GitHub↗

    Thai Word Segmentation and Part-of-Speech Tagging with Deep Learning

    Python
    Vezi pe GitHub↗41
  • ksopyla/awesome-nlp-polishAvatar ksopyla

    ksopyla/awesome-nlp-polish

    308Vezi pe GitHub↗

    A curated list of resources dedicated to Natural Language Processing (NLP) in polish. Models, tools, datasets.

    Vezi pe GitHub↗308
  • mikahama/uralicnlpAvatar mikahama

    mikahama/uralicNLP

    98Vezi pe GitHub↗

    Natural language processing for many languages

    Python
    Vezi pe GitHub↗98
  • narimann2/parsianalyzerAvatar NarimanN2

    NarimanN2/ParsiAnalyzer

    166Vezi pe GitHub↗

    Persian Analyzer for Elasticsearch.

    Java
    Vezi pe GitHub↗166
  • phuonglh/vn.vitkAvatar phuonglh

    phuonglh/vn.vitk

    218Vezi pe GitHub↗

    NOTE: This repos is now obsolete. Interested programmers should consider to use the new repo vlp (github.com/phuonglh/vlp) We have preferred using Scala instead of Java since 2016.

    Java
    Vezi pe GitHub↗218