awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
wb14123 avatar

wb14123/couplet-dataset

0
View on GitHub↗
745 stars·222 forks·Python·AGPL-3.0·13 views

Couplet Dataset

Dataset for couplets. 70万条对联数据库。

Features

  • Natural Language Processing - Listed in the “Natural Language Processing” section of the FunNLP awesome list.
  • Text Generation - Dataset for training couplet generation models.
  • Natural Language Corpora - Large-scale dataset of traditional Chinese couplets.

Star history

Star history chart for wb14123/couplet-datasetStar history chart for wb14123/couplet-dataset

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Couplet Dataset

These projects share indexed features with Couplet Dataset. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • complementizer/wcep-mds-datasetcomplementizer avatar

    complementizer/wcep-mds-dataset

    61View on GitHub↗

    The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

    Python
    View on GitHub↗61
  • dbamman/litbankdbamman avatar

    dbamman/litbank

    377View on GitHub↗

    Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.

    Python
    View on GitHub↗377
  • asyml/texarasyml avatar

    asyml/texar

    2,392View on GitHub↗

    Toolkit for Machine Learning, Natural Language Processing, and Text Generation, in TensorFlow. This is part of the CASL project: http://casl-project.ai/

    Pythonbertcasl-projectdata-processing
    View on GitHub↗2,392
  • didi/chinesenlpdidi avatar

    didi/ChineseNLP

    1,811View on GitHub↗

    Datasets, SOTA results of every fields of Chinese NLP

    HTMLchinese-nlpchinese-word-segmentationentity-linking
    View on GitHub↗1,811
Compare all 30 related projects→

Frequently asked questions

What does wb14123/couplet-dataset do?

Dataset for couplets. 70万条对联数据库。

What are the main features of wb14123/couplet-dataset?

The main features of wb14123/couplet-dataset are: Natural Language Processing, Text Generation, Natural Language Corpora.

Which projects share features with wb14123/couplet-dataset?

Projects with overlapping indexed features include: dbamman/litbank — Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational… edinburghnlp/opus-100-corpus — OPUS-100. asyml/texar — Toolkit for Machine Learning, Natural Language Processing, and Text Generation, in TensorFlow. This is part of the… complementizer/wcep-mds-dataset — The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events,… didi/chinesenlp — Datasets, SOTA results of every fields of Chinese NLP. embedding/chinese-word-vectors — This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It…