awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
didi avatar

didi/ChineseNLPArchived

0
View on GitHub↗
1,811 stars·260 forks·HTML·7 viewschinesenlp.xyz↗

ChineseNLP

Datasets, SOTA results of every fields of Chinese NLP

Features

  • Natural Language Processing - Listed in the “Natural Language Processing” section of the FunNLP awesome list.
  • Natural Language Corpora - Open tasks and benchmarks for Chinese NLP research.

Star history

Star history chart for didi/chinesenlpStar history chart for didi/chinesenlp

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does didi/chinesenlp do?

Datasets, SOTA results of every fields of Chinese NLP

What are the main features of didi/chinesenlp?

The main features of didi/chinesenlp are: Natural Language Processing, Natural Language Corpora.

What are some open-source alternatives to didi/chinesenlp?

Open-source alternatives to didi/chinesenlp include: edinburghnlp/opus-100-corpus — OPUS-100. facebookresearch/laser — LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text… complementizer/wcep-mds-dataset — The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events,… dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It… google-research-datasets/dakshina — The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each…

Open-source alternatives to ChineseNLP

Similar open-source projects, ranked by how many features they share with ChineseNLP.
  • dbamman/litbankdbamman avatar

    dbamman/litbank

    377View on GitHub↗

    Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.

    Python
    View on GitHub↗377
  • edinburghnlp/opus-100-corpusEdinburghNLP avatar

    EdinburghNLP/opus-100-corpus

    93View on GitHub↗

    OPUS-100

    Python
    View on GitHub↗93
  • complementizer/wcep-mds-datasetcomplementizer avatar

    complementizer/wcep-mds-dataset

    61View on GitHub↗

    The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

    Python
    View on GitHub↗61
  • embedding/chinese-word-vectorsEmbedding avatar

    Embedding/Chinese-Word-Vectors

    12,227View on GitHub↗

    This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu

    Pythonchinesechinese-word-segmentationembedding
    View on GitHub↗12,227
See all 30 alternatives to ChineseNLP→