How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…
The main features of complementizer/wcep-mds-dataset are: Natural Language Processing, Natural Language Corpora.
Projects with overlapping indexed features include: edinburghnlp/opus-100-corpus — OPUS-100. facebookresearch/laser — LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text… dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… didi/chinesenlp — Datasets, SOTA results of every fields of Chinese NLP. embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It… google-research-datasets/dakshina — The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each…
Datasets, SOTA results of every fields of Chinese NLP
Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.
This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu