awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
complementizer avatar

complementizer/wcep-mds-dataset

0
View on GitHub↗
61 stars·15 forks·Python·MIT·4 vues

Wcep Mds Dataset

The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

Features

  • Natural Language Processing - Listed in the “Natural Language Processing” section of the FunNLP awesome list.
  • Natural Language Corpora - Dataset for multi-document summarization tasks.

Historique des stars

Graphique de l'historique des stars pour complementizer/wcep-mds-datasetGraphique de l'historique des stars pour complementizer/wcep-mds-dataset

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à Wcep Mds Dataset

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Wcep Mds Dataset.
  • didi/chinesenlpAvatar de didi

    didi/ChineseNLP

    1,811Voir sur GitHub↗

    Datasets, SOTA results of every fields of Chinese NLP

    HTMLchinese-nlpchinese-word-segmentationentity-linking
    Voir sur GitHub↗1,811
  • edinburghnlp/opus-100-corpusAvatar de EdinburghNLP

    EdinburghNLP/opus-100-corpus

    93Voir sur GitHub↗

    OPUS-100

    Python
    Voir sur GitHub↗93
  • dbamman/litbankAvatar de dbamman

    dbamman/litbank

    377Voir sur GitHub↗

    Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.

    Python
    Voir sur GitHub↗377
  • embedding/chinese-word-vectorsAvatar de Embedding

    Embedding/Chinese-Word-Vectors

    12,227Voir sur GitHub↗

    This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu

    Pythonchinesechinese-word-segmentationembedding
    Voir sur GitHub↗12,227
Voir les 30 alternatives à Wcep Mds Dataset→

Questions fréquentes

Que fait complementizer/wcep-mds-dataset ?

The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

Quelles sont les fonctionnalités principales de complementizer/wcep-mds-dataset ?

Les fonctionnalités principales de complementizer/wcep-mds-dataset sont : Natural Language Processing, Natural Language Corpora.

Quelles sont les alternatives open-source à complementizer/wcep-mds-dataset ?

Les alternatives open-source à complementizer/wcep-mds-dataset incluent : edinburghnlp/opus-100-corpus — OPUS-100. facebookresearch/laser — LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text… dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… didi/chinesenlp — Datasets, SOTA results of every fields of Chinese NLP. embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It… google-research-datasets/dakshina — The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each…