awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to blmoistawinde/harvesttext

Open-source alternatives to HarvestText

30 open-source projects similar to blmoistawinde/harvesttext, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best HarvestText alternative.

  • liuhuanyong/wordmultisensedisambiguationliuhuanyong 的头像

    liuhuanyong/WordMultiSenseDisambiguation

    131在 GitHub 上查看↗

    WordMultiSenseDisambiguation, chinese multi-wordsense disambiguation based on online bake knowledge base and semantic embedding similarity compute,基于百科知识库的中文词语多词义/义项获取与特定句子词语语义消歧.

    Python
    在 GitHub 上查看↗131
  • skishore/makemeahanziskishore 的头像

    skishore/makemeahanzi

    2,535在 GitHub 上查看↗

    Free, open-source Chinese character data

    JavaScript
    在 GitHub 上查看↗2,535
  • jacksonllee/pycantonesejacksonllee 的头像

    jacksonllee/pycantonese

    409在 GitHub 上查看↗

    Cantonese Linguistics and NLP

    Pythoncantonesecomputational-linguisticsjyutping
    在 GitHub 上查看↗409
  • liuhuanyong/domainwordsdictliuhuanyong 的头像

    liuhuanyong/DomainWordsDict

    769在 GitHub 上查看↗

    DomainWordsDict, Chinese words dict that contains more than 68 domains, which can be used as text classification、knowledge enhance task。涵盖68个领域、共计916万词的专业词典知识库,可用于文本分类、知识增强、领域词汇库扩充等自然语言处理应用。

    在 GitHub 上查看↗769
  • qingyujean/sscqingyujean 的头像

    qingyujean/ssc

    226在 GitHub 上查看↗

    基于“音形码”的中文字符串相似度计算方法

    Python
    在 GitHub 上查看↗226

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • tinyfool/chinesewithenglishtinyfool 的头像

    tinyfool/ChineseWithEnglish

    52在 GitHub 上查看↗

    绝对有趣的中文发音引擎 funny chinese text to speech enginee

    在 GitHub 上查看↗52
  • fighting41love/coconlpfighting41love 的头像

    fighting41love/cocoNLP

    1,130在 GitHub 上查看↗

    A Chinese information extraction tool.

    Python
    在 GitHub 上查看↗1,130
  • huggingface/tokenizershuggingface 的头像

    huggingface/tokenizers

    10,825在 GitHub 上查看↗

    This project is a high-performance library for converting raw text into tokens and IDs for machine learning models. It functions as a fast text encoder and a text preprocessing pipeline designed to transform strings into numerical representations with high throughput for research and production. The library includes a subword tokenizer trainer used to analyze text datasets and create custom vocabularies using algorithms such as byte-pair encoding and wordpiece. It provides capabilities for subword vocabulary training and text alignment, allowing character offsets to be tracked during normaliz

    Rustbertgptlanguage-model
    在 GitHub 上查看↗10,825
  • keredson/wordninjakeredson 的头像

    keredson/wordninja

    872在 GitHub 上查看↗

    Probabilistically split concatenated words using NLP based on English Wikipedia unigram frequencies.

    Python
    在 GitHub 上查看↗872
  • liuhuanyong/crimekgassitantliuhuanyong 的头像

    liuhuanyong/CrimeKgAssitant

    1,580在 GitHub 上查看↗

    Crime assistant including crime type prediction and crime consult service based on nlp methods and crime kg,罪名法务智能项目,内容包括856项罪名知识图谱, 基于280万罪名训练库的罪名预测,基于20W法务问答对的13类问题分类与法律资讯问答功能.

    Python
    在 GitHub 上查看↗1,580
  • observerss/textfilterobserverss 的头像

    observerss/textfilter

    2,113在 GitHub 上查看↗

    敏感词过滤的几种实现+某1w词敏感词库

    Python
    在 GitHub 上查看↗2,113
  • pwxcoo/chinese-xinhuapwxcoo 的头像

    pwxcoo/chinese-xinhua

    11,572在 GitHub 上查看↗

    Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona

    Pythonchinesechinese-characterschinese-language
    在 GitHub 上查看↗11,572
  • skydark/nstoolsskydark 的头像

    skydark/nstools

    678在 GitHub 上查看↗

    Some meaningless nscripter tools.

    Python
    在 GitHub 上查看↗678
  • wainshine/company-names-corpuswainshine 的头像

    wainshine/Company-Names-Corpus

    1,293在 GitHub 上查看↗

    公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。

    companycorpusdataset
    在 GitHub 上查看↗1,293
  • 1eez/1039761eez 的头像

    1eez/103976

    1,034在 GitHub 上查看↗

    103976个英语单词库(sql版,csv版,Excel版)包含英文单词,中文翻译,单词的词性及多种词义,执行SQL语句就可以生成表,支持SQL Server,MySQL等多种数据库

    PLpgSQL
    在 GitHub 上查看↗1,034
  • berniey/hanziconvberniey 的头像

    berniey/hanziconv

    190在 GitHub 上查看↗

    Hanzi Converter for Traditional and Simplified Chinese

    Python
    在 GitHub 上查看↗190
  • glassywing/bi-lstm-crfGlassyWing 的头像

    GlassyWing/bi-lstm-crf

    384在 GitHub 上查看↗

    使用keras实现的基于Bi-LSTM CRF的中文分词+词性标注

    Python
    在 GitHub 上查看↗384
  • glassywing/transformer-word-segmenterGlassyWing 的头像

    GlassyWing/transformer-word-segmenter

    163在 GitHub 上查看↗

    Sequence labeling base on universal transformer (Transformer encoder) and CRF; 基于Universal Transformer CRF 的中文分词和词性标注

    Python
    在 GitHub 上查看↗163
  • jeongukjae/python-mecabjeongukjae 的头像

    jeongukjae/python-mecab

    28在 GitHub 上查看↗

    A repository to bind mecab for Python 3.5+. Not using swig nor pybind. (Not Maintained Now)

    C++mecabpython-c-extensiontext-preprocessing
    在 GitHub 上查看↗28
  • kaleidophon/token2indexKaleidophon 的头像

    Kaleidophon/token2index

    50在 GitHub 上查看↗

    A lightweight but powerful library to build token indices for NLP tasks, compatible with major Deep Learning frameworks like PyTorch and Tensorflow.

    Pythondeep-learningdeeplearningi2t
    在 GitHub 上查看↗50
  • kfcd/chaizikfcd 的头像

    kfcd/chaizi

    811在 GitHub 上查看↗

    漢語拆字字典

    chinesechinese-characterscomponents
    在 GitHub 上查看↗811
  • kyubyong/g2pcKyubyong 的头像

    Kyubyong/g2pC

    245在 GitHub 上查看↗

    g2pC: A Context-aware Grapheme-to-Phoneme Conversion module for Chinese

    Pythonchinese-nlpchinese-word-segmentationcrf
    在 GitHub 上查看↗245
  • mozillazg/phrase-pinyin-datamozillazg 的头像

    mozillazg/phrase-pinyin-data

    530在 GitHub 上查看↗

    词语拼音数据

    Pythonpinyinpinyin-data
    在 GitHub 上查看↗530
  • mozillazg/python-pinyinmozillazg 的头像

    mozillazg/python-pinyin

    5,325在 GitHub 上查看↗

    python-pinyin is a Python library for transliterating simplified and traditional Chinese characters into phonetic pinyin. It functions as a transliteration system that converts text while supporting tone sandhi and providing utilities to transform pinyin between different formats, such as numeric tones, accent marks, or phonetic initials. The library features a polyphonic character resolver that analyzes surrounding word context to select the correct pronunciation for characters with multiple sounds. It also includes a customizable dictionary system that allows the extension of default transl

    Pythonchinesehanzihanzi-pinyin
    在 GitHub 上查看↗5,325
  • opennmt/tokenizerOpenNMT 的头像

    OpenNMT/Tokenizer

    333在 GitHub 上查看↗

    Fast and customizable text tokenization library with BPE and SentencePiece support

    C++bpecppicu
    在 GitHub 上查看↗333
  • philipperemy/name-datasetphilipperemy 的头像

    philipperemy/name-dataset

    1,002在 GitHub 上查看↗

    The Python library for names.

    Pythondatasetnamenamed-entity-recognition
    在 GitHub 上查看↗1,002
  • rainarch/sentibridgerainarch 的头像

    rainarch/SentiBridge

    639在 GitHub 上查看↗

    SentiBridge: A Knowledge Base for Entity-Sentiment Representation

    Pythonknowledge-graphsentiment-analysis
    在 GitHub 上查看↗639
  • artificiai/multilingual-latent-dirichlet-allocation-ldaArtificiAI 的头像

    ArtificiAI/Multilingual-Latent-Dirichlet-Allocation-LDA

    83在 GitHub 上查看↗

    A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

    Pythonclusteringenglishfrench
    在 GitHub 上查看↗83
  • alibaba-edu/simple-effective-text-matching-pytorchA

    alibaba-edu/simple-effective-text-matching-pytorch

    0在 GitHub 上查看↗
    在 GitHub 上查看↗0
  • artidoro/qloraartidoro 的头像

    artidoro/qlora

    10,929在 GitHub 上查看↗

    This project is a quantized fine-tuning framework for large language models. It implements a low-rank adaptation library and a four-bit quantizer to reduce the GPU memory requirements needed to train large models. The framework utilizes four-bit quantization and low-rank adapters to enable model training on consumer-grade hardware. It further reduces the memory footprint through double quantization and a paged optimizer that offloads states to system RAM. The system supports distributed training across multiple GPUs to handle larger parameter scales and includes utilities for custom dataset

    Jupyter Notebook
    在 GitHub 上查看↗10,929