How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.
Datasets, SOTA results of every fields of Chinese NLP
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…
非常全的文言文(古文)-现代文平行语料
The main features of niutrans/classical-modern are: Natural Language Processing, Natural Language Corpora.
Open-source alternatives to niutrans/classical-modern include: didi/chinesenlp — Datasets, SOTA results of every fields of Chinese NLP. embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It… complementizer/wcep-mds-dataset — The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events,… dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… edinburghnlp/opus-100-corpus — OPUS-100. facebookresearch/laser — LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text…