awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to rainarch/sentibridge

Open-source alternatives to SentiBridge

30 open-source projects similar to rainarch/sentibridge, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best SentiBridge alternative.

  • observerss/textfilterالصورة الرمزية لـ observerss

    observerss/textfilter

    2,113عرض على GitHub↗

    敏感词过滤的几种实现+某1w词敏感词库

    Python
    عرض على GitHub↗2,113
  • pwxcoo/chinese-xinhuaالصورة الرمزية لـ pwxcoo

    pwxcoo/chinese-xinhua

    11,572عرض على GitHub↗

    Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona

    Pythonchinesechinese-characterschinese-language
    عرض على GitHub↗11,572
  • wainshine/company-names-corpusالصورة الرمزية لـ wainshine

    wainshine/Company-Names-Corpus

    1,293عرض على GitHub↗

    公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。

    companycorpusdataset
    عرض على GitHub↗1,293
  • kfcd/chaiziالصورة الرمزية لـ kfcd

    kfcd/chaizi

    811عرض على GitHub↗

    漢語拆字字典

    chinesechinese-characterscomponents
    عرض على GitHub↗811
  • liuhuanyong/domainwordsdictالصورة الرمزية لـ liuhuanyong

    liuhuanyong/DomainWordsDict

    769عرض على GitHub↗

    DomainWordsDict, Chinese words dict that contains more than 68 domains, which can be used as text classification、knowledge enhance task。涵盖68个领域、共计916万词的专业词典知识库,可用于文本分类、知识增强、领域词汇库扩充等自然语言处理应用。

    عرض على GitHub↗769

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • skydark/nstoolsالصورة الرمزية لـ skydark

    skydark/nstools

    678عرض على GitHub↗

    Some meaningless nscripter tools.

    Python
    عرض على GitHub↗678
  • sophonplus/chinesenlpcorpusالصورة الرمزية لـ SophonPlus

    SophonPlus/ChineseNlpCorpus

    6,568عرض على GitHub↗

    搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。

    Jupyter Notebook
    عرض على GitHub↗6,568
  • huyingxi/synonymsالصورة الرمزية لـ huyingxi

    huyingxi/Synonyms

    5,107عرض على GitHub↗

    Synonyms is a Chinese natural language processing tool focused on semantic analysis. It provides capabilities for Chinese word segmentation, part-of-speech tagging, and the retrieval of synonyms based on semantic proximity. The project converts words and sentences into numerical vector representations to calculate similarity scores. This allows for the determination of semantic proximity between different phrases and the identification of chatbot intent through sentence comparison. The system also includes tools for automated keyword extraction and importance ranking to identify significant

    Python
    عرض على GitHub↗5,107
  • keredson/wordninjaالصورة الرمزية لـ keredson

    keredson/wordninja

    872عرض على GitHub↗

    Probabilistically split concatenated words using NLP based on English Wikipedia unigram frequencies.

    Python
    عرض على GitHub↗872
  • kaleidophon/token2indexالصورة الرمزية لـ Kaleidophon

    Kaleidophon/token2index

    50عرض على GitHub↗

    A lightweight but powerful library to build token indices for NLP tasks, compatible with major Deep Learning frameworks like PyTorch and Tensorflow.

    Pythondeep-learningdeeplearningi2t
    عرض على GitHub↗50
  • kyubyong/g2pcالصورة الرمزية لـ Kyubyong

    Kyubyong/g2pC

    245عرض على GitHub↗

    g2pC: A Context-aware Grapheme-to-Phoneme Conversion module for Chinese

    Pythonchinese-nlpchinese-word-segmentationcrf
    عرض على GitHub↗245
  • mozillazg/python-pinyinالصورة الرمزية لـ mozillazg

    mozillazg/python-pinyin

    5,325عرض على GitHub↗

    python-pinyin is a Python library for transliterating simplified and traditional Chinese characters into phonetic pinyin. It functions as a transliteration system that converts text while supporting tone sandhi and providing utilities to transform pinyin between different formats, such as numeric tones, accent marks, or phonetic initials. The library features a polyphonic character resolver that analyzes surrounding word context to select the correct pronunciation for characters with multiple sounds. It also includes a customizable dictionary system that allows the extension of default transl

    Pythonchinesehanzihanzi-pinyin
    عرض على GitHub↗5,325
  • opennmt/tokenizerالصورة الرمزية لـ OpenNMT

    OpenNMT/Tokenizer

    333عرض على GitHub↗

    Fast and customizable text tokenization library with BPE and SentencePiece support

    C++bpecppicu
    عرض على GitHub↗333
  • skishore/makemeahanziالصورة الرمزية لـ skishore

    skishore/makemeahanzi

    2,535عرض على GitHub↗

    Free, open-source Chinese character data

    JavaScript
    عرض على GitHub↗2,535
  • chinese-poetry/chinese-poetryالصورة الرمزية لـ chinese-poetry

    chinese-poetry/chinese-poetry

    51,906عرض على GitHub↗

    This project is a comprehensive dataset and archive of classical Chinese poetry, prose, and Confucian classics. It serves as a digital humanities corpus, providing machine-readable access to hundreds of thousands of poems and detailed poet biographies, specifically spanning the Tang and Song dynasties. The collection is distinguished by its scholarly depth, incorporating textual variation annotations to track disputed characters across different source editions. It also includes tonal pattern mapping to describe the rhythmic and phonetic structures of the verse, alongside a popularity ranking

    JavaScriptchinesechinese-poetryci
    عرض على GitHub↗51,906
  • fighting41love/coconlpالصورة الرمزية لـ fighting41love

    fighting41love/cocoNLP

    1,130عرض على GitHub↗

    A Chinese information extraction tool.

    Python
    عرض على GitHub↗1,130
  • berniey/hanziconvالصورة الرمزية لـ berniey

    berniey/hanziconv

    190عرض على GitHub↗

    Hanzi Converter for Traditional and Simplified Chinese

    Python
    عرض على GitHub↗190
  • embedding/chinese-word-vectorsالصورة الرمزية لـ Embedding

    Embedding/Chinese-Word-Vectors

    12,227عرض على GitHub↗

    This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu

    Pythonchinesechinese-word-segmentationembedding
    عرض على GitHub↗12,227
  • glassywing/transformer-word-segmenterالصورة الرمزية لـ GlassyWing

    GlassyWing/transformer-word-segmenter

    163عرض على GitHub↗

    Sequence labeling base on universal transformer (Transformer encoder) and CRF; 基于Universal Transformer CRF 的中文分词和词性标注

    Python
    عرض على GitHub↗163
  • huggingface/tokenizersالصورة الرمزية لـ huggingface

    huggingface/tokenizers

    10,825عرض على GitHub↗

    This project is a high-performance library for converting raw text into tokens and IDs for machine learning models. It functions as a fast text encoder and a text preprocessing pipeline designed to transform strings into numerical representations with high throughput for research and production. The library includes a subword tokenizer trainer used to analyze text datasets and create custom vocabularies using algorithms such as byte-pair encoding and wordpiece. It provides capabilities for subword vocabulary training and text alignment, allowing character offsets to be tracked during normaliz

    Rustbertgptlanguage-model
    عرض على GitHub↗10,825
  • jacksonllee/pycantoneseالصورة الرمزية لـ jacksonllee

    jacksonllee/pycantonese

    409عرض على GitHub↗

    Cantonese Linguistics and NLP

    Pythoncantonesecomputational-linguisticsjyutping
    عرض على GitHub↗409
  • jeongukjae/python-mecabالصورة الرمزية لـ jeongukjae

    jeongukjae/python-mecab

    28عرض على GitHub↗

    A repository to bind mecab for Python 3.5+. Not using swig nor pybind. (Not Maintained Now)

    C++mecabpython-c-extensiontext-preprocessing
    عرض على GitHub↗28
  • 1eez/103976الصورة الرمزية لـ 1eez

    1eez/103976

    1,034عرض على GitHub↗

    103976个英语单词库(sql版,csv版,Excel版)包含英文单词,中文翻译,单词的词性及多种词义,执行SQL语句就可以生成表,支持SQL Server,MySQL等多种数据库

    PLpgSQL
    عرض على GitHub↗1,034
  • liuhuanyong/crimekgassitantالصورة الرمزية لـ liuhuanyong

    liuhuanyong/CrimeKgAssitant

    1,580عرض على GitHub↗

    Crime assistant including crime type prediction and crime consult service based on nlp methods and crime kg,罪名法务智能项目,内容包括856项罪名知识图谱, 基于280万罪名训练库的罪名预测,基于20W法务问答对的13类问题分类与法律资讯问答功能.

    Python
    عرض على GitHub↗1,580
  • liuhuanyong/wordmultisensedisambiguationالصورة الرمزية لـ liuhuanyong

    liuhuanyong/WordMultiSenseDisambiguation

    131عرض على GitHub↗

    WordMultiSenseDisambiguation, chinese multi-wordsense disambiguation based on online bake knowledge base and semantic embedding similarity compute,基于百科知识库的中文词语多词义/义项获取与特定句子词语语义消歧.

    Python
    عرض على GitHub↗131
  • mozillazg/phrase-pinyin-dataالصورة الرمزية لـ mozillazg

    mozillazg/phrase-pinyin-data

    530عرض على GitHub↗

    词语拼音数据

    Pythonpinyinpinyin-data
    عرض على GitHub↗530
  • blmoistawinde/harvesttextالصورة الرمزية لـ blmoistawinde

    blmoistawinde/HarvestText

    2,621عرض على GitHub↗

    文本挖掘和预处理工具(文本清洗、新词发现、情感分析、实体识别链接、关键词抽取、知识抽取、句法分析等),无监督或弱监督方法

    Pythondependency-parsergiteeharvesttext
    عرض على GitHub↗2,621
  • philipperemy/name-datasetالصورة الرمزية لـ philipperemy

    philipperemy/name-dataset

    1,002عرض على GitHub↗

    The Python library for names.

    Pythondatasetnamenamed-entity-recognition
    عرض على GitHub↗1,002
  • qingyujean/sscالصورة الرمزية لـ qingyujean

    qingyujean/ssc

    226عرض على GitHub↗

    基于“音形码”的中文字符串相似度计算方法

    Python
    عرض على GitHub↗226
  • glassywing/bi-lstm-crfالصورة الرمزية لـ GlassyWing

    GlassyWing/bi-lstm-crf

    384عرض على GitHub↗

    使用keras实现的基于Bi-LSTM CRF的中文分词+词性标注

    Python
    عرض على GitHub↗384