awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to embedding/chinese-word-vectors

Projects sharing features with Chinese Word Vectors

30 open-source projects similar to embedding/chinese-word-vectors, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • ymcui/chinese-bert-wwmymcui avatar

    ymcui/Chinese-BERT-wwm

    10,212View on GitHub↗

    Chinese-BERT-wwm is a pre-trained transformer model and encoder designed for Chinese natural language processing. It converts Chinese text into dense vector representations to be used across various natural language processing applications. The model utilizes a whole word masking strategy during pre-training, masking entire words rather than individual characters. This approach is designed to improve the capture of semantic meaning and language structure within Chinese datasets. The project covers a range of downstream tasks including text classification, sequence labeling, and reading compr

    Pythonbertbert-wwmbert-wwm-ext
    View on GitHub↗10,212
  • thunlp/openclapthunlp avatar

    thunlp/OpenCLaP

    984View on GitHub↗

    Open Chinese Language Pre-trained Model Zoo

    View on GitHub↗984
  • zhangyics/chinese-abbreviation-datasetzhangyics avatar

    zhangyics/Chinese-abbreviation-dataset

    199View on GitHub↗

    This is a corpus of Chinese abbreviation, including negative full forms.

    View on GitHub↗199
  • wainshine/chinese-names-corpuswainshine avatar

    wainshine/Chinese-Names-Corpus

    4,303View on GitHub↗

    This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages. The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources. The available data su

    corpusdatasetdict
    View on GitHub↗4,303

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • sophonplus/chinesenlpcorpusSophonPlus avatar

    SophonPlus/ChineseNlpCorpus

    6,568View on GitHub↗

    搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。

    Jupyter Notebook
    View on GitHub↗6,568
  • hit-scir/ltpHIT-SCIR avatar

    HIT-SCIR/ltp

    5,253View on GitHub↗

    This is a Chinese natural language processing toolkit providing a suite of tools for word segmentation, part-of-speech tagging, and named entity recognition. It includes a neural dependency parser for analyzing syntactic and semantic relationships between words and a machine learning training suite for creating custom linguistic models using annotated datasets. The toolkit distinguishes itself through its deployment flexibility, offering a dockerized server and a web service interface that exposes processing capabilities via API. It supports the use of pretrained models and allows for the int

    Pythonchinese-nlpmachine-learningnatural-language-processing
    View on GitHub↗5,253
  • lancopku/pkuseg-pythonlancopku avatar

    lancopku/pkuseg-python

    6,707View on GitHub↗

    pkuseg-python is a Chinese word segmentation toolkit and natural language processing library. It provides specialized models for splitting Chinese text into words across various domains, including news, medical, and web content, and includes a tool for assigning grammatical parts of speech tags to segmented words. The library allows for the training of custom segmentation models using annotated datasets and supports the integration of user-defined dictionaries to ensure specialized terminology is recognized correctly. It employs a multi-threaded execution engine to process large volumes of Ch

    Python
    View on GitHub↗6,707
  • isnowfy/snownlpisnowfy avatar

    isnowfy/snownlp

    6,631View on GitHub↗

    SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation, sentiment analysis, document classification, and phonetic transliteration. The library includes capabilities for training and saving custom machine learning models for tokenization and sentiment analysis using raw training datasets. It covers a range of linguistic processing areas, including parts of speech tagging, sentence splitting, and text similarity measurement. The toolkit also provides utilities for extracting key information through text summarization and calculating word im

    Python
    View on GitHub↗6,631
  • huggingface/sentence-transformershuggingface avatar

    huggingface/sentence-transformers

    18,817View on GitHub↗

    This project is a transformer-based framework for generating dense and sparse vector embeddings of text and multimodal data. It serves as a library for fine-tuning models to perform semantic similarity tasks, retrieval, and reranking. The system is distinguished by its support for diverse architectural patterns, including bi-encoders for fast similarity search and cross-encoders for high-precision reranking. It provides dedicated pipelines for multimodal embeddings, mapping text and images into a shared vector space, and implements knowledge distillation to compress large models into smaller,

    Python
    View on GitHub↗18,817
  • thunlp-aipoet/datasetsTHUNLP-AIPoet avatar

    THUNLP-AIPoet/Datasets

    239View on GitHub↗

    Poetry-related datasets developed by THUAIPoet (Jiuge) group.

    chinesecorpuspoetry-generation
    View on GitHub↗239
  • didi/chinesenlpdidi avatar

    didi/ChineseNLP

    1,811View on GitHub↗

    Datasets, SOTA results of every fields of Chinese NLP

    HTMLchinese-nlpchinese-word-segmentationentity-linking
    View on GitHub↗1,811
  • mattdesl/dictionary-of-colour-combinationsmattdesl avatar

    mattdesl/dictionary-of-colour-combinations

    464View on GitHub↗

    palettes from A Dictionary of Colour Combinations

    Python
    View on GitHub↗464
  • openlmlab/gaokao-benchOpenLMLab avatar

    OpenLMLab/GAOKAO-Bench

    760View on GitHub↗

    GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.

    Python
    View on GitHub↗760
  • observerss/textfilterobserverss avatar

    observerss/textfilter

    2,113View on GitHub↗

    敏感词过滤的几种实现+某1w词敏感词库

    Python
    View on GitHub↗2,113
  • niutrans/classical-modernNiuTrans avatar

    NiuTrans/Classical-Modern

    1,450View on GitHub↗

    非常全的文言文(古文)-现代文平行语料

    Pythoncorpusparallel-corpustraditional-and-simplified-chinese
    View on GitHub↗1,450
  • oye93/chinese-nlp-corpusOYE93 avatar

    OYE93/Chinese-NLP-Corpus

    922View on GitHub↗

    Collections of Chinese NLP corpus

    Pythonchinese-nlpcorpusdatasets
    View on GitHub↗922
  • liuhuanyong/chineseembeddingliuhuanyong avatar

    liuhuanyong/ChineseEmbedding

    454View on GitHub↗

    Chinese Embedding collection incling token ,postag ,pinyin,dependency,word embedding.中文自然语言处理向量合集,包括字向量,拼音向量,词向量,词性向量,依存关系向量.共5种类型的向量

    Python
    View on GitHub↗454
  • google-research-datasets/dakshinagoogle-research-datasets avatar

    google-research-datasets/dakshina

    212View on GitHub↗

    The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each language, the dataset includes a large collection of native script Wikipedia text, a romanization lexicon of words in the native script with attested romanizations, and some full sentence parallel data in both a native script of the language and the basic Latin alphabet.

    View on GitHub↗212
  • chinese-poetry/chinese-poetrychinese-poetry avatar

    chinese-poetry/chinese-poetry

    51,906View on GitHub↗

    This project is a comprehensive dataset and archive of classical Chinese poetry, prose, and Confucian classics. It serves as a digital humanities corpus, providing machine-readable access to hundreds of thousands of poems and detailed poet biographies, specifically spanning the Tang and Song dynasties. The collection is distinguished by its scholarly depth, incorporating textual variation annotations to track disputed characters across different source editions. It also includes tonal pattern mapping to describe the rhythmic and phonetic structures of the verse, alongside a popularity ranking

    JavaScriptchinesechinese-poetryci
    View on GitHub↗51,906
  • khiajohnson/spice-corpuskhiajohnson avatar

    khiajohnson/SpiCE-Corpus

    40View on GitHub↗

    An open-access corpus of conversational bilingual speech in Cantonese and English

    JavaScriptbilingual-corporacantonese-languagecorpus
    View on GitHub↗40
  • complementizer/wcep-mds-datasetcomplementizer avatar

    complementizer/wcep-mds-dataset

    61View on GitHub↗

    The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

    Python
    View on GitHub↗61
  • several27/fakenewscorpusseveral27 avatar

    several27/FakeNewsCorpus

    413View on GitHub↗

    A dataset of millions of news articles scraped from a curated list of data sources.

    artificial-intelligencecorpusdatabase
    View on GitHub↗413
  • google-research/bertgoogle-research avatar

    google-research/bert

    39,869View on GitHub↗

    This project is a transformer-based language model and natural language processing toolkit designed to generate deep contextual representations of text. By utilizing a transformer-based encoder architecture, the system processes input sequences through stacked self-attention layers to capture the semantic meaning of tokens based on their surrounding sentence structure. The model distinguishes itself through bidirectional contextual processing, which analyzes text in both directions simultaneously, and masked language modeling, which trains the system by predicting hidden tokens within a seque

    Pythongooglenatural-language-processingnatural-language-understanding
    View on GitHub↗39,869
  • pwxcoo/chinese-xinhuapwxcoo avatar

    pwxcoo/chinese-xinhua

    11,572View on GitHub↗

    Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona

    Pythonchinesechinese-characterschinese-language
    View on GitHub↗11,572
  • paddlepaddle/erniePaddlePaddle avatar

    PaddlePaddle/ERNIE

    7,717View on GitHub↗

    ERNIE is a development toolkit for training, fine-tuning, and deploying large language models built on the PaddlePaddle deep learning platform. It provides a comprehensive suite of core components, including an inference server for vision and language models, a training and fine-tuning toolkit, and a framework for building retrieval-augmented generation systems using private knowledge bases. The project features multimodal AI models capable of reasoning across text, images, and video to perform complex visual understanding and information extraction. It distinguishes itself through specialize

    Pythonernieernie-45ernie-45-vl
    View on GitHub↗7,717
  • dbamman/litbankdbamman avatar

    dbamman/litbank

    377View on GitHub↗

    Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.

    Python
    View on GitHub↗377
  • edinburghnlp/opus-100-corpusEdinburghNLP avatar

    EdinburghNLP/opus-100-corpus

    93View on GitHub↗

    OPUS-100

    Python
    View on GitHub↗93
  • kfcd/chaizikfcd avatar

    kfcd/chaizi

    811View on GitHub↗

    漢語拆字字典

    chinesechinese-characterscomponents
    View on GitHub↗811
  • facebookresearch/laserfacebookresearch avatar

    facebookresearch/LASER

    3,659View on GitHub↗

    LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text mining tool that maps sentences from multiple languages into a shared vector space for similarity and classification tasks. The system converts raw text into fixed-length embeddings, enabling the discovery of translation pairs by calculating the vector distance between sentences. This shared representation allows for cross-lingual document classification, where a model trained on one language can be used to categorize documents in another. The library includes a sentence-piece t

    Jupyter Notebook
    View on GitHub↗3,659
  • panhaiqi/ancientpoetrypanhaiqi avatar

    panhaiqi/AncientPoetry

    137View on GitHub↗

    古诗词语料库

    View on GitHub↗137