How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.
Datasets, SOTA results of every fields of Chinese NLP
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…
The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each language, the dataset includes a large collection of native script Wikipedia text, a romanization lexicon of words in the native script with attested romanizations, and some full sentence parallel data in both a native script of the language and the basic Latin alphabet.
The main features of google-research-datasets/dakshina are: Natural Language Processing, Natural Language Corpora.
Projects with overlapping indexed features include: didi/chinesenlp — Datasets, SOTA results of every fields of Chinese NLP. embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It… complementizer/wcep-mds-dataset — The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events,… dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… edinburghnlp/opus-100-corpus — OPUS-100. facebookresearch/laser — LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text…