awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to wb14123/couplet-dataset

Open-source alternatives to Couplet Dataset

30 open-source projects similar to wb14123/couplet-dataset, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Couplet Dataset alternative.

  • thunlp/openclapthunlp avatar

    thunlp/OpenCLaP

    984View on GitHub↗

    Open Chinese Language Pre-trained Model Zoo

    View on GitHub↗984
  • yc9701/pansoriyc9701 avatar

    yc9701/pansori

    139View on GitHub↗

    Tools for ASR Corpus Generation from Online Video

    Pythoncorpusdata-pipelinedataset-generation
    View on GitHub↗139
  • oye93/chinese-nlp-corpusOYE93 avatar

    OYE93/Chinese-NLP-Corpus

    922View on GitHub↗

    Collections of Chinese NLP corpus

    Pythonchinese-nlpcorpusdatasets
    View on GitHub↗922
  • openlmlab/gaokao-benchOpenLMLab avatar

    OpenLMLab/GAOKAO-Bench

    760View on GitHub↗

    GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.

    Python
    View on GitHub↗760
  • xiangyuecn/areacity-jsspider-statsgovxiangyuecn avatar

    xiangyuecn/AreaCity-JsSpider-StatsGov

    6,672View on GitHub↗

    This project is an administrative GIS toolset that provides a comprehensive dataset of China's administrative divisions, including provinces, cities, districts, and townships. It functions as a coordinate system transformer and a boundary converter for transforming geographic data into standard formats. The toolset distinguishes itself through the ability to convert administrative boundary data between CSV, GeoJSON, Shapefiles, and SQL. It includes specialized utilities for coordinate system transformation between GCJ-02, BD-09, WGS-84, and CGCS2000 standards to ensure accuracy across differe

    JavaScript
    View on GitHub↗6,672

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • ymcui/chinese-bert-wwmymcui avatar

    ymcui/Chinese-BERT-wwm

    10,212View on GitHub↗

    Chinese-BERT-wwm is a pre-trained transformer model and encoder designed for Chinese natural language processing. It converts Chinese text into dense vector representations to be used across various natural language processing applications. The model utilizes a whole word masking strategy during pre-training, masking entire words rather than individual characters. This approach is designed to improve the capture of semantic meaning and language structure within Chinese datasets. The project covers a range of downstream tasks including text classification, sequence labeling, and reading compr

    Pythonbertbert-wwmbert-wwm-ext
    View on GitHub↗10,212
  • dbamman/litbankdbamman avatar

    dbamman/litbank

    377View on GitHub↗

    Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.

    Python
    View on GitHub↗377
  • embedding/chinese-word-vectorsEmbedding avatar

    Embedding/Chinese-Word-Vectors

    12,227View on GitHub↗

    This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu

    Pythonchinesechinese-word-segmentationembedding
    View on GitHub↗12,227
  • khiajohnson/spice-corpuskhiajohnson avatar

    khiajohnson/SpiCE-Corpus

    40View on GitHub↗

    An open-access corpus of conversational bilingual speech in Cantonese and English

    JavaScriptbilingual-corporacantonese-languagecorpus
    View on GitHub↗40
  • niutrans/classical-modernNiuTrans avatar

    NiuTrans/Classical-Modern

    1,450View on GitHub↗

    非常全的文言文(古文)-现代文平行语料

    Pythoncorpusparallel-corpustraditional-and-simplified-chinese
    View on GitHub↗1,450
  • sophonplus/chinesenlpcorpusSophonPlus avatar

    SophonPlus/ChineseNlpCorpus

    6,568View on GitHub↗

    搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。

    Jupyter Notebook
    View on GitHub↗6,568
  • wainshine/chinese-names-corpuswainshine avatar

    wainshine/Chinese-Names-Corpus

    4,303View on GitHub↗

    This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages. The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources. The available data su

    corpusdatasetdict
    View on GitHub↗4,303
  • ymcui/chinese-rc-datasetsymcui avatar

    ymcui/Chinese-RC-Datasets

    221View on GitHub↗

    Collections of Chinese reading comprehension datasets

    question-answeringreading-comprehension
    View on GitHub↗221
  • zhangyics/chinese-abbreviation-datasetzhangyics avatar

    zhangyics/Chinese-abbreviation-dataset

    199View on GitHub↗

    This is a corpus of Chinese abbreviation, including negative full forms.

    View on GitHub↗199
  • asyml/texarasyml avatar

    asyml/texar

    2,392View on GitHub↗

    Toolkit for Machine Learning, Natural Language Processing, and Text Generation, in TensorFlow. This is part of the CASL project: http://casl-project.ai/

    Pythonbertcasl-projectdata-processing
    View on GitHub↗2,392
  • complementizer/wcep-mds-datasetcomplementizer avatar

    complementizer/wcep-mds-dataset

    61View on GitHub↗

    The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

    Python
    View on GitHub↗61
  • didi/chinesenlpdidi avatar

    didi/ChineseNLP

    1,811View on GitHub↗

    Datasets, SOTA results of every fields of Chinese NLP

    HTMLchinese-nlpchinese-word-segmentationentity-linking
    View on GitHub↗1,811
  • edinburghnlp/opus-100-corpusEdinburghNLP avatar

    EdinburghNLP/opus-100-corpus

    93View on GitHub↗

    OPUS-100

    Python
    View on GitHub↗93
  • facebookresearch/laserfacebookresearch avatar

    facebookresearch/LASER

    3,659View on GitHub↗

    LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text mining tool that maps sentences from multiple languages into a shared vector space for similarity and classification tasks. The system converts raw text into fixed-length embeddings, enabling the discovery of translation pairs by calculating the vector distance between sentences. This shared representation allows for cross-lingual document classification, where a model trained on one language can be used to categorize documents in another. The library includes a sentence-piece t

    Jupyter Notebook
    View on GitHub↗3,659
  • google-research-datasets/dakshinagoogle-research-datasets avatar

    google-research-datasets/dakshina

    212View on GitHub↗

    The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each language, the dataset includes a large collection of native script Wikipedia text, a romanization lexicon of words in the native script with attested romanizations, and some full sentence parallel data in both a native script of the language and the basic Latin alphabet.

    View on GitHub↗212
  • liuhuanyong/chineseembeddingliuhuanyong avatar

    liuhuanyong/ChineseEmbedding

    454View on GitHub↗

    Chinese Embedding collection incling token ,postag ,pinyin,dependency,word embedding.中文自然语言处理向量合集,包括字向量,拼音向量,词向量,词性向量,依存关系向量.共5种类型的向量

    Python
    View on GitHub↗454
  • mattdesl/dictionary-of-colour-combinationsmattdesl avatar

    mattdesl/dictionary-of-colour-combinations

    464View on GitHub↗

    palettes from A Dictionary of Colour Combinations

    Python
    View on GitHub↗464
  • panhaiqi/ancientpoetrypanhaiqi avatar

    panhaiqi/AncientPoetry

    137View on GitHub↗

    古诗词语料库

    View on GitHub↗137
  • several27/fakenewscorpusseveral27 avatar

    several27/FakeNewsCorpus

    413View on GitHub↗

    A dataset of millions of news articles scraped from a curated list of data sources.

    artificial-intelligencecorpusdatabase
    View on GitHub↗413
  • thunlp-aipoet/datasetsTHUNLP-AIPoet avatar

    THUNLP-AIPoet/Datasets

    239View on GitHub↗

    Poetry-related datasets developed by THUAIPoet (Jiuge) group.

    chinesecorpuspoetry-generation
    View on GitHub↗239
  • facebookresearch/flow_matchingfacebookresearch avatar

    facebookresearch/flow_matching

    4,562View on GitHub↗

    This project is a PyTorch-based generative model framework designed to transform noise into complex data distributions by learning vector fields and probability paths. It serves as a multimodal generative toolkit for producing synthetic text and images through learned probability flows. The library distinguishes itself by supporting continuous, discrete, and Riemannian manifold integrations. This allows the framework to handle a variety of data types, including categorical data via discrete-state flow matching and non-Euclidean spaces through Riemannian manifold integration. The toolkit cove

    Python
    View on GitHub↗4,562
  • artificiai/multilingual-latent-dirichlet-allocation-ldaArtificiAI avatar

    ArtificiAI/Multilingual-Latent-Dirichlet-Allocation-LDA

    83View on GitHub↗

    A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

    Pythonclusteringenglishfrench
    View on GitHub↗83
  • alibaba-edu/simple-effective-text-matching-pytorchA

    alibaba-edu/simple-effective-text-matching-pytorch

    0View on GitHub↗
    View on GitHub↗0
  • artidoro/qloraartidoro avatar

    artidoro/qlora

    10,929View on GitHub↗

    This project is a quantized fine-tuning framework for large language models. It implements a low-rank adaptation library and a four-bit quantizer to reduce the GPU memory requirements needed to train large models. The framework utilizes four-bit quantization and low-rank adapters to enable model training on consumer-grade hardware. It further reduces the memory footprint through double quantization and a paged optimizer that offloads states to system RAM. The system supports distributed training across multiple GPUs to handle larger parameter scales and includes utilities for custom dataset

    Jupyter Notebook
    View on GitHub↗10,929
  • arongdari/topic-model-lecture-notearongdari avatar

    arongdari/topic-model-lecture-note

    22View on GitHub↗

    lecture notes for probabilistic topic models using ipython notebook

    View on GitHub↗22