awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to blmoistawinde/harvesttext

Open-source alternatives to HarvestText

30 open-source projects similar to blmoistawinde/harvesttext, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best HarvestText alternative.

  • liuhuanyong/wordmultisensedisambiguationliuhuanyong का अवतार

    liuhuanyong/WordMultiSenseDisambiguation

    131GitHub पर देखें↗

    WordMultiSenseDisambiguation, chinese multi-wordsense disambiguation based on online bake knowledge base and semantic embedding similarity compute,基于百科知识库的中文词语多词义/义项获取与特定句子词语语义消歧.

    Python
    GitHub पर देखें↗131
  • skishore/makemeahanziskishore का अवतार

    skishore/makemeahanzi

    2,535GitHub पर देखें↗

    Free, open-source Chinese character data

    JavaScript
    GitHub पर देखें↗2,535
  • jacksonllee/pycantonesejacksonllee का अवतार

    jacksonllee/pycantonese

    409GitHub पर देखें↗

    Cantonese Linguistics and NLP

    Pythoncantonesecomputational-linguisticsjyutping
    GitHub पर देखें↗409
  • liuhuanyong/domainwordsdictliuhuanyong का अवतार

    liuhuanyong/DomainWordsDict

    769GitHub पर देखें↗

    DomainWordsDict, Chinese words dict that contains more than 68 domains, which can be used as text classification、knowledge enhance task。涵盖68个领域、共计916万词的专业词典知识库,可用于文本分类、知识增强、领域词汇库扩充等自然语言处理应用。

    GitHub पर देखें↗769
  • qingyujean/sscqingyujean का अवतार

    qingyujean/ssc

    226GitHub पर देखें↗

    基于“音形码”的中文字符串相似度计算方法

    Python
    GitHub पर देखें↗226

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Find more with AI search
  • tinyfool/chinesewithenglishtinyfool का अवतार

    tinyfool/ChineseWithEnglish

    52GitHub पर देखें↗

    绝对有趣的中文发音引擎 funny chinese text to speech enginee

    GitHub पर देखें↗52
  • fighting41love/coconlpfighting41love का अवतार

    fighting41love/cocoNLP

    1,130GitHub पर देखें↗

    A Chinese information extraction tool.

    Python
    GitHub पर देखें↗1,130
  • huggingface/tokenizershuggingface का अवतार

    huggingface/tokenizers

    10,825GitHub पर देखें↗

    This project is a high-performance library for converting raw text into tokens and IDs for machine learning models. It functions as a fast text encoder and a text preprocessing pipeline designed to transform strings into numerical representations with high throughput for research and production. The library includes a subword tokenizer trainer used to analyze text datasets and create custom vocabularies using algorithms such as byte-pair encoding and wordpiece. It provides capabilities for subword vocabulary training and text alignment, allowing character offsets to be tracked during normaliz

    Rustbertgptlanguage-model
    GitHub पर देखें↗10,825
  • keredson/wordninjakeredson का अवतार

    keredson/wordninja

    872GitHub पर देखें↗

    Probabilistically split concatenated words using NLP based on English Wikipedia unigram frequencies.

    Python
    GitHub पर देखें↗872
  • liuhuanyong/crimekgassitantliuhuanyong का अवतार

    liuhuanyong/CrimeKgAssitant

    1,580GitHub पर देखें↗

    Crime assistant including crime type prediction and crime consult service based on nlp methods and crime kg,罪名法务智能项目,内容包括856项罪名知识图谱, 基于280万罪名训练库的罪名预测,基于20W法务问答对的13类问题分类与法律资讯问答功能.

    Python
    GitHub पर देखें↗1,580
  • observerss/textfilterobserverss का अवतार

    observerss/textfilter

    2,113GitHub पर देखें↗

    敏感词过滤的几种实现+某1w词敏感词库

    Python
    GitHub पर देखें↗2,113
  • pwxcoo/chinese-xinhuapwxcoo का अवतार

    pwxcoo/chinese-xinhua

    11,572GitHub पर देखें↗

    Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona

    Pythonchinesechinese-characterschinese-language
    GitHub पर देखें↗11,572
  • skydark/nstoolsskydark का अवतार

    skydark/nstools

    678GitHub पर देखें↗

    Some meaningless nscripter tools.

    Python
    GitHub पर देखें↗678
  • wainshine/company-names-corpuswainshine का अवतार

    wainshine/Company-Names-Corpus

    1,293GitHub पर देखें↗

    公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。

    companycorpusdataset
    GitHub पर देखें↗1,293
  • 1eez/1039761eez का अवतार

    1eez/103976

    1,034GitHub पर देखें↗

    103976个英语单词库(sql版,csv版,Excel版)包含英文单词,中文翻译,单词的词性及多种词义,执行SQL语句就可以生成表,支持SQL Server,MySQL等多种数据库

    PLpgSQL
    GitHub पर देखें↗1,034
  • berniey/hanziconvberniey का अवतार

    berniey/hanziconv

    190GitHub पर देखें↗

    Hanzi Converter for Traditional and Simplified Chinese

    Python
    GitHub पर देखें↗190
  • glassywing/bi-lstm-crfGlassyWing का अवतार

    GlassyWing/bi-lstm-crf

    384GitHub पर देखें↗

    使用keras实现的基于Bi-LSTM CRF的中文分词+词性标注

    Python
    GitHub पर देखें↗384
  • glassywing/transformer-word-segmenterGlassyWing का अवतार

    GlassyWing/transformer-word-segmenter

    163GitHub पर देखें↗

    Sequence labeling base on universal transformer (Transformer encoder) and CRF; 基于Universal Transformer CRF 的中文分词和词性标注

    Python
    GitHub पर देखें↗163
  • jeongukjae/python-mecabjeongukjae का अवतार

    jeongukjae/python-mecab

    28GitHub पर देखें↗

    A repository to bind mecab for Python 3.5+. Not using swig nor pybind. (Not Maintained Now)

    C++mecabpython-c-extensiontext-preprocessing
    GitHub पर देखें↗28
  • kaleidophon/token2indexKaleidophon का अवतार

    Kaleidophon/token2index

    50GitHub पर देखें↗

    A lightweight but powerful library to build token indices for NLP tasks, compatible with major Deep Learning frameworks like PyTorch and Tensorflow.

    Pythondeep-learningdeeplearningi2t
    GitHub पर देखें↗50
  • kfcd/chaizikfcd का अवतार

    kfcd/chaizi

    811GitHub पर देखें↗

    漢語拆字字典

    chinesechinese-characterscomponents
    GitHub पर देखें↗811
  • kyubyong/g2pcKyubyong का अवतार

    Kyubyong/g2pC

    245GitHub पर देखें↗

    g2pC: A Context-aware Grapheme-to-Phoneme Conversion module for Chinese

    Pythonchinese-nlpchinese-word-segmentationcrf
    GitHub पर देखें↗245
  • mozillazg/phrase-pinyin-datamozillazg का अवतार

    mozillazg/phrase-pinyin-data

    530GitHub पर देखें↗

    词语拼音数据

    Pythonpinyinpinyin-data
    GitHub पर देखें↗530
  • mozillazg/python-pinyinmozillazg का अवतार

    mozillazg/python-pinyin

    5,325GitHub पर देखें↗

    python-pinyin is a Python library for transliterating simplified and traditional Chinese characters into phonetic pinyin. It functions as a transliteration system that converts text while supporting tone sandhi and providing utilities to transform pinyin between different formats, such as numeric tones, accent marks, or phonetic initials. The library features a polyphonic character resolver that analyzes surrounding word context to select the correct pronunciation for characters with multiple sounds. It also includes a customizable dictionary system that allows the extension of default transl

    Pythonchinesehanzihanzi-pinyin
    GitHub पर देखें↗5,325
  • opennmt/tokenizerOpenNMT का अवतार

    OpenNMT/Tokenizer

    333GitHub पर देखें↗

    Fast and customizable text tokenization library with BPE and SentencePiece support

    C++bpecppicu
    GitHub पर देखें↗333
  • philipperemy/name-datasetphilipperemy का अवतार

    philipperemy/name-dataset

    1,002GitHub पर देखें↗

    The Python library for names.

    Pythondatasetnamenamed-entity-recognition
    GitHub पर देखें↗1,002
  • rainarch/sentibridgerainarch का अवतार

    rainarch/SentiBridge

    639GitHub पर देखें↗

    SentiBridge: A Knowledge Base for Entity-Sentiment Representation

    Pythonknowledge-graphsentiment-analysis
    GitHub पर देखें↗639
  • artificiai/multilingual-latent-dirichlet-allocation-ldaArtificiAI का अवतार

    ArtificiAI/Multilingual-Latent-Dirichlet-Allocation-LDA

    83GitHub पर देखें↗

    A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

    Pythonclusteringenglishfrench
    GitHub पर देखें↗83
  • alibaba-edu/simple-effective-text-matching-pytorchA

    alibaba-edu/simple-effective-text-matching-pytorch

    0GitHub पर देखें↗
    GitHub पर देखें↗0
  • artidoro/qloraartidoro का अवतार

    artidoro/qlora

    10,929GitHub पर देखें↗

    This project is a quantized fine-tuning framework for large language models. It implements a low-rank adaptation library and a four-bit quantizer to reduce the GPU memory requirements needed to train large models. The framework utilizes four-bit quantization and low-rank adapters to enable model training on consumer-grade hardware. It further reduces the memory footprint through double quantization and a paged optimizer that offloads states to system RAM. The system supports distributed training across multiple GPUs to handle larger parameter scales and includes utilities for custom dataset

    Jupyter Notebook
    GitHub पर देखें↗10,929