How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
Datasets, SOTA results of every fields of Chinese NLP
The main features of didi/chinesenlp are: Natural Language Processing, Natural Language Corpora.
Open-source alternatives to didi/chinesenlp include: edinburghnlp/opus-100-corpus — OPUS-100. facebookresearch/laser — LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text… complementizer/wcep-mds-dataset — The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events,… dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It… google-research-datasets/dakshina — The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each…
Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…
This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu