awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Embedding avatar

Embedding/Chinese-Word-Vectors

0
View on GitHub↗
12,227 stars·2,325 forks·Python·Apache-2.0·13 vues

Chinese Word Vectors

This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models.

The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision.

The project covers semantic vector search and the development of Chinese natural language processing pipelines. It also includes tools for text embedding evaluation, utilizing analogy-based quality tests to measure morphological and semantic accuracy.

Features

  • Chinese - Provides various pre-trained vector sets trained on different Chinese corpora to capture diverse linguistic patterns.
  • Chinese NLP Libraries - Offers pre-trained embeddings specifically designed to improve semantic understanding for Chinese text NLP pipelines.
  • Chinese Word Embedding Collections - Provides a set of pre-trained dense and sparse word vectors trained on diverse Chinese corpora.
  • N-Gram Co-occurrence Models - Captures semantic relationships by analyzing the frequency of adjacent word fragments and character sequences during training.
  • Pre-trained Vector Repositories - Maintains a collection of dense and sparse word embeddings derived from large-scale diverse Chinese text corpora.
  • Hybrid Sparse-Dense Embeddings - Implements a hybrid of dense semantic vectors and sparse keyword-based vectors for balanced efficiency and precision.
  • Pre-trained Vector Libraries - Provides a collection of linguistic representations used to improve the accuracy of semantic and morphological analysis.
  • Word Embedding Datasets - Provides access to word embedding datasets trained on diverse corpora for NLP tasks.
  • Embedding - Evaluates vector performance using specialized toolkits and datasets focused on morphological and semantic analogies.
  • Analogy Solvers - Provides tools to solve morphological and semantic word relationship puzzles via vector arithmetic.
  • Vector Search - Enables finding words or phrases with similar meanings using high-dimensional vector distance.
  • Corpus-Specific Vectorizations - Generates distinct vector spaces based on the specific linguistic characteristics of different training data sources.
  • Natural Language Processing - Listed in the “Natural Language Processing” section of the FunNLP awesome list.
  • Pretrained Models and Embeddings - Large collection of pretrained Chinese word embeddings.
  • Research and Datasets - Large-scale pre-trained word embeddings for Chinese NLP tasks.
  • Corpus and Datasets - Collection of pre-trained Chinese word embeddings.
  • Natural Language Corpora - Pre-trained word vector representations for Chinese.

Historique des stars

Graphique de l'historique des stars pour embedding/chinese-word-vectorsGraphique de l'historique des stars pour embedding/chinese-word-vectors

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Questions fréquentes

Que fait embedding/chinese-word-vectors ?

This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models.

Quelles sont les fonctionnalités principales de embedding/chinese-word-vectors ?

Les fonctionnalités principales de embedding/chinese-word-vectors sont : Chinese, Chinese NLP Libraries, Chinese Word Embedding Collections, N-Gram Co-occurrence Models, Pre-trained Vector Repositories, Hybrid Sparse-Dense Embeddings, Pre-trained Vector Libraries, Word Embedding Datasets.

Quelles sont les alternatives open-source à embedding/chinese-word-vectors ?

Les alternatives open-source à embedding/chinese-word-vectors incluent : ymcui/chinese-bert-wwm — Chinese-BERT-wwm is a pre-trained transformer model and encoder designed for Chinese natural language processing. It… sophonplus/chinesenlpcorpus — 搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。. wainshine/chinese-names-corpus — This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis… thunlp/openclap — Open Chinese Language Pre-trained Model Zoo. zhangyics/chinese-abbreviation-dataset — This is a corpus of Chinese abbreviation, including negative full forms. hit-scir/ltp — This is a Chinese natural language processing toolkit providing a suite of tools for word segmentation, part-of-speech…

Alternatives open source à Chinese Word Vectors

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Chinese Word Vectors.
  • ymcui/chinese-bert-wwmAvatar de ymcui

    ymcui/Chinese-BERT-wwm

    10,212Voir sur GitHub↗

    Chinese-BERT-wwm is a pre-trained transformer model and encoder designed for Chinese natural language processing. It converts Chinese text into dense vector representations to be used across various natural language processing applications. The model utilizes a whole word masking strategy during pre-training, masking entire words rather than individual characters. This approach is designed to improve the capture of semantic meaning and language structure within Chinese datasets. The project covers a range of downstream tasks including text classification, sequence labeling, and reading compr

    Pythonbertbert-wwmbert-wwm-ext
    Voir sur GitHub↗10,212
  • thunlp/openclapAvatar de thunlp

    thunlp/OpenCLaP

    984Voir sur GitHub↗

    Open Chinese Language Pre-trained Model Zoo

    Voir sur GitHub↗984
  • sophonplus/chinesenlpcorpusAvatar de SophonPlus

    SophonPlus/ChineseNlpCorpus

    6,568Voir sur GitHub↗

    搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。

    Jupyter Notebook
    Voir sur GitHub↗6,568
  • wainshine/chinese-names-corpusAvatar de wainshine

    wainshine/Chinese-Names-Corpus

    4,303Voir sur GitHub↗

    This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages. The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources. The available data su

    corpusdatasetdict
    Voir sur GitHub↗4,303
Voir les 30 alternatives à Chinese Word Vectors→