How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…
Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.
Toolkit for Machine Learning, Natural Language Processing, and Text Generation, in TensorFlow. This is part of the CASL project: http://casl-project.ai/
Datasets, SOTA results of every fields of Chinese NLP
Dataset for couplets. 70万条对联数据库。
The main features of wb14123/couplet-dataset are: Natural Language Processing, Text Generation, Natural Language Corpora.
Projects with overlapping indexed features include: dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… edinburghnlp/opus-100-corpus — OPUS-100. asyml/texar — Toolkit for Machine Learning, Natural Language Processing, and Text Generation, in TensorFlow. This is part of the… complementizer/wcep-mds-dataset — The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events,… didi/chinesenlp — Datasets, SOTA results of every fields of Chinese NLP. embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It…