Datasets, SOTA results of every fields of Chinese NLP
Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.
This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…
Les fonctionnalités principales de complementizer/wcep-mds-dataset sont : Natural Language Processing, Natural Language Corpora.
Les alternatives open-source à complementizer/wcep-mds-dataset incluent : edinburghnlp/opus-100-corpus — OPUS-100. facebookresearch/laser — LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text… dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… didi/chinesenlp — Datasets, SOTA results of every fields of Chinese NLP. embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It… google-research-datasets/dakshina — The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each…