awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to thunlp/openclap

Projects sharing features with OpenCLaP

30 open-source projects similar to thunlp/openclap, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • ymcui/chinese-bert-wwmymcui avatar

    ymcui/Chinese-BERT-wwm

    10,212View on GitHub↗

    Chinese-BERT-wwm is a pre-trained transformer model and encoder designed for Chinese natural language processing. It converts Chinese text into dense vector representations to be used across various natural language processing applications. The model utilizes a whole word masking strategy during pre-training, masking entire words rather than individual characters. This approach is designed to improve the capture of semantic meaning and language structure within Chinese datasets. The project covers a range of downstream tasks including text classification, sequence labeling, and reading compr

    Pythonbertbert-wwmbert-wwm-ext
    View on GitHub↗10,212
  • wainshine/chinese-names-corpuswainshine avatar

    wainshine/Chinese-Names-Corpus

    4,303View on GitHub↗

    This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages. The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources. The available data su

    corpusdatasetdict
    View on GitHub↗4,303
  • embedding/chinese-word-vectorsEmbedding avatar

    Embedding/Chinese-Word-Vectors

    12,227View on GitHub↗

    This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu

    Pythonchinesechinese-word-segmentationembedding
    View on GitHub↗12,227

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • zhangyics/chinese-abbreviation-datasetzhangyics avatar

    zhangyics/Chinese-abbreviation-dataset

    199View on GitHub↗

    This is a corpus of Chinese abbreviation, including negative full forms.

    View on GitHub↗199
  • sophonplus/chinesenlpcorpusSophonPlus avatar

    SophonPlus/ChineseNlpCorpus

    6,568View on GitHub↗

    搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。

    Jupyter Notebook
    View on GitHub↗6,568
  • wb14123/couplet-datasetwb14123 avatar

    wb14123/couplet-dataset

    745View on GitHub↗

    Dataset for couplets. 70万条对联数据库。

    Pythondataset
    View on GitHub↗745
  • thunlp-aipoet/datasetsTHUNLP-AIPoet avatar

    THUNLP-AIPoet/Datasets

    239View on GitHub↗

    Poetry-related datasets developed by THUAIPoet (Jiuge) group.

    chinesecorpuspoetry-generation
    View on GitHub↗239
  • ymcui/chinese-rc-datasetsymcui avatar

    ymcui/Chinese-RC-Datasets

    221View on GitHub↗

    Collections of Chinese reading comprehension datasets

    question-answeringreading-comprehension
    View on GitHub↗221
  • niutrans/classical-modernNiuTrans avatar

    NiuTrans/Classical-Modern

    1,450View on GitHub↗

    非常全的文言文(古文)-现代文平行语料

    Pythoncorpusparallel-corpustraditional-and-simplified-chinese
    View on GitHub↗1,450
  • khiajohnson/spice-corpuskhiajohnson avatar

    khiajohnson/SpiCE-Corpus

    40View on GitHub↗

    An open-access corpus of conversational bilingual speech in Cantonese and English

    JavaScriptbilingual-corporacantonese-languagecorpus
    View on GitHub↗40
  • pwxcoo/chinese-xinhuapwxcoo avatar

    pwxcoo/chinese-xinhua

    11,572View on GitHub↗

    Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona

    Pythonchinesechinese-characterschinese-language
    View on GitHub↗11,572
  • rainarch/sentibridgerainarch avatar

    rainarch/SentiBridge

    639View on GitHub↗

    SentiBridge: A Knowledge Base for Entity-Sentiment Representation

    Pythonknowledge-graphsentiment-analysis
    View on GitHub↗639
  • wainshine/company-names-corpuswainshine avatar

    wainshine/Company-Names-Corpus

    1,293View on GitHub↗

    公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。

    companycorpusdataset
    View on GitHub↗1,293
  • yc9701/pansoriyc9701 avatar

    yc9701/pansori

    139View on GitHub↗

    Tools for ASR Corpus Generation from Online Video

    Pythoncorpusdata-pipelinedataset-generation
    View on GitHub↗139
  • observerss/textfilterobserverss avatar

    observerss/textfilter

    2,113View on GitHub↗

    敏感词过滤的几种实现+某1w词敏感词库

    Python
    View on GitHub↗2,113
  • edinburghnlp/opus-100-corpusEdinburghNLP avatar

    EdinburghNLP/opus-100-corpus

    93View on GitHub↗

    OPUS-100

    Python
    View on GitHub↗93
  • didi/chinesenlpdidi avatar

    didi/ChineseNLP

    1,811View on GitHub↗

    Datasets, SOTA results of every fields of Chinese NLP

    HTMLchinese-nlpchinese-word-segmentationentity-linking
    View on GitHub↗1,811
  • complementizer/wcep-mds-datasetcomplementizer avatar

    complementizer/wcep-mds-dataset

    61View on GitHub↗

    The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

    Python
    View on GitHub↗61
  • google-research-datasets/dakshinagoogle-research-datasets avatar

    google-research-datasets/dakshina

    212View on GitHub↗

    The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each language, the dataset includes a large collection of native script Wikipedia text, a romanization lexicon of words in the native script with attested romanizations, and some full sentence parallel data in both a native script of the language and the basic Latin alphabet.

    View on GitHub↗212
  • kfcd/chaizikfcd avatar

    kfcd/chaizi

    811View on GitHub↗

    漢語拆字字典

    chinesechinese-characterscomponents
    View on GitHub↗811
  • liuhuanyong/chineseembeddingliuhuanyong avatar

    liuhuanyong/ChineseEmbedding

    454View on GitHub↗

    Chinese Embedding collection incling token ,postag ,pinyin,dependency,word embedding.中文自然语言处理向量合集,包括字向量,拼音向量,词向量,词性向量,依存关系向量.共5种类型的向量

    Python
    View on GitHub↗454
  • mattdesl/dictionary-of-colour-combinationsmattdesl avatar

    mattdesl/dictionary-of-colour-combinations

    464View on GitHub↗

    palettes from A Dictionary of Colour Combinations

    Python
    View on GitHub↗464
  • oye93/chinese-nlp-corpusOYE93 avatar

    OYE93/Chinese-NLP-Corpus

    922View on GitHub↗

    Collections of Chinese NLP corpus

    Pythonchinese-nlpcorpusdatasets
    View on GitHub↗922
  • panhaiqi/ancientpoetrypanhaiqi avatar

    panhaiqi/AncientPoetry

    137View on GitHub↗

    古诗词语料库

    View on GitHub↗137
  • chinese-poetry/chinese-poetrychinese-poetry avatar

    chinese-poetry/chinese-poetry

    51,906View on GitHub↗

    This project is a comprehensive dataset and archive of classical Chinese poetry, prose, and Confucian classics. It serves as a digital humanities corpus, providing machine-readable access to hundreds of thousands of poems and detailed poet biographies, specifically spanning the Tang and Song dynasties. The collection is distinguished by its scholarly depth, incorporating textual variation annotations to track disputed characters across different source editions. It also includes tonal pattern mapping to describe the rhythmic and phonetic structures of the verse, alongside a popularity ranking

    JavaScriptchinesechinese-poetryci
    View on GitHub↗51,906
  • several27/fakenewscorpusseveral27 avatar

    several27/FakeNewsCorpus

    413View on GitHub↗

    A dataset of millions of news articles scraped from a curated list of data sources.

    artificial-intelligencecorpusdatabase
    View on GitHub↗413
  • dbamman/litbankdbamman avatar

    dbamman/litbank

    377View on GitHub↗

    Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.

    Python
    View on GitHub↗377
  • facebookresearch/laserfacebookresearch avatar

    facebookresearch/LASER

    3,659View on GitHub↗

    LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text mining tool that maps sentences from multiple languages into a shared vector space for similarity and classification tasks. The system converts raw text into fixed-length embeddings, enabling the discovery of translation pairs by calculating the vector distance between sentences. This shared representation allows for cross-lingual document classification, where a model trained on one language can be used to categorize documents in another. The library includes a sentence-piece t

    Jupyter Notebook
    View on GitHub↗3,659
  • xiangyuecn/areacity-jsspider-statsgovxiangyuecn avatar

    xiangyuecn/AreaCity-JsSpider-StatsGov

    6,672View on GitHub↗

    This project is an administrative GIS toolset that provides a comprehensive dataset of China's administrative divisions, including provinces, cities, districts, and townships. It functions as a coordinate system transformer and a boundary converter for transforming geographic data into standard formats. The toolset distinguishes itself through the ability to convert administrative boundary data between CSV, GeoJSON, Shapefiles, and SQL. It includes specialized utilities for coordinate system transformation between GCJ-02, BD-09, WGS-84, and CGCS2000 standards to ensure accuracy across differe

    JavaScript
    View on GitHub↗6,672
  • openlmlab/gaokao-benchOpenLMLab avatar

    OpenLMLab/GAOKAO-Bench

    760View on GitHub↗

    GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.

    Python
    View on GitHub↗760