awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to thunlp/openclap

Open-source alternatives to OpenCLaP

30 open-source projects similar to thunlp/openclap, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best OpenCLaP alternative.

  • ymcui/chinese-bert-wwmAvatar de ymcui

    ymcui/Chinese-BERT-wwm

    10,212Voir sur GitHub↗

    Chinese-BERT-wwm is a pre-trained transformer model and encoder designed for Chinese natural language processing. It converts Chinese text into dense vector representations to be used across various natural language processing applications. The model utilizes a whole word masking strategy during pre-training, masking entire words rather than individual characters. This approach is designed to improve the capture of semantic meaning and language structure within Chinese datasets. The project covers a range of downstream tasks including text classification, sequence labeling, and reading compr

    Pythonbertbert-wwmbert-wwm-ext
    Voir sur GitHub↗10,212
  • wainshine/chinese-names-corpusAvatar de wainshine

    wainshine/Chinese-Names-Corpus

    4,303Voir sur GitHub↗

    This project is a curated collection of Chinese names, surnames, and kinship terms designed for linguistic analysis and natural language processing. It functions as a multilingual name dataset and a training resource for named entity recognition, providing a unified repository of names across Chinese, Japanese, and English languages. The project includes a synthetic name generator that creates realistic person names by applying analyzed naming patterns and demographic data. It also provides a cleaned Chinese idiom lexicon gathered and deduplicated from multiple sources. The available data su

    corpusdatasetdict
    Voir sur GitHub↗4,303
  • embedding/chinese-word-vectorsAvatar de Embedding

    Embedding/Chinese-Word-Vectors

    12,227Voir sur GitHub↗

    This project is a collection of pre-trained dense and sparse word vectors trained on diverse Chinese corpora. It serves as a library of linguistic representations and an NLP vector dataset designed to improve the accuracy of semantic and morphological analysis in text models. The collection provides corpus-specific representations and utilizes n-gram co-occurrence modeling to capture diverse linguistic patterns. It includes a hybrid of dense-sparse vectors to balance computational efficiency and semantic precision. The project covers semantic vector search and the development of Chinese natu

    Pythonchinesechinese-word-segmentationembedding
    Voir sur GitHub↗12,227

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Find more with AI search
  • zhangyics/chinese-abbreviation-datasetAvatar de zhangyics

    zhangyics/Chinese-abbreviation-dataset

    199Voir sur GitHub↗

    This is a corpus of Chinese abbreviation, including negative full forms.

    Voir sur GitHub↗199
  • sophonplus/chinesenlpcorpusAvatar de SophonPlus

    SophonPlus/ChineseNlpCorpus

    6,568Voir sur GitHub↗

    搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。

    Jupyter Notebook
    Voir sur GitHub↗6,568
  • wb14123/couplet-datasetAvatar de wb14123

    wb14123/couplet-dataset

    745Voir sur GitHub↗

    Dataset for couplets. 70万条对联数据库。

    Pythondataset
    Voir sur GitHub↗745
  • thunlp-aipoet/datasetsAvatar de THUNLP-AIPoet

    THUNLP-AIPoet/Datasets

    239Voir sur GitHub↗

    Poetry-related datasets developed by THUAIPoet (Jiuge) group.

    chinesecorpuspoetry-generation
    Voir sur GitHub↗239
  • ymcui/chinese-rc-datasetsAvatar de ymcui

    ymcui/Chinese-RC-Datasets

    221Voir sur GitHub↗

    Collections of Chinese reading comprehension datasets

    question-answeringreading-comprehension
    Voir sur GitHub↗221
  • niutrans/classical-modernAvatar de NiuTrans

    NiuTrans/Classical-Modern

    1,450Voir sur GitHub↗

    非常全的文言文(古文)-现代文平行语料

    Pythoncorpusparallel-corpustraditional-and-simplified-chinese
    Voir sur GitHub↗1,450
  • khiajohnson/spice-corpusAvatar de khiajohnson

    khiajohnson/SpiCE-Corpus

    40Voir sur GitHub↗

    An open-access corpus of conversational bilingual speech in Cantonese and English

    JavaScriptbilingual-corporacantonese-languagecorpus
    Voir sur GitHub↗40
  • pwxcoo/chinese-xinhuaAvatar de pwxcoo

    pwxcoo/chinese-xinhua

    11,572Voir sur GitHub↗

    Chinese-xinhua is an open-source repository providing a comprehensive, machine-readable collection of Chinese linguistic data. It serves as a structured archive of dictionary entries, idioms, and phrases designed for programmatic access and integration into language processing applications. The project organizes complex linguistic information into consistent, schema-driven object structures that facilitate rapid lookups and data portability. By utilizing key-value indexing and structured text serialization, the dataset enables developers to implement advanced natural language search functiona

    Pythonchinesechinese-characterschinese-language
    Voir sur GitHub↗11,572
  • rainarch/sentibridgeAvatar de rainarch

    rainarch/SentiBridge

    639Voir sur GitHub↗

    SentiBridge: A Knowledge Base for Entity-Sentiment Representation

    Pythonknowledge-graphsentiment-analysis
    Voir sur GitHub↗639
  • wainshine/company-names-corpusAvatar de wainshine

    wainshine/Company-Names-Corpus

    1,293Voir sur GitHub↗

    公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。

    companycorpusdataset
    Voir sur GitHub↗1,293
  • yc9701/pansoriAvatar de yc9701

    yc9701/pansori

    139Voir sur GitHub↗

    Tools for ASR Corpus Generation from Online Video

    Pythoncorpusdata-pipelinedataset-generation
    Voir sur GitHub↗139
  • observerss/textfilterAvatar de observerss

    observerss/textfilter

    2,113Voir sur GitHub↗

    敏感词过滤的几种实现+某1w词敏感词库

    Python
    Voir sur GitHub↗2,113
  • edinburghnlp/opus-100-corpusAvatar de EdinburghNLP

    EdinburghNLP/opus-100-corpus

    93Voir sur GitHub↗

    OPUS-100

    Python
    Voir sur GitHub↗93
  • didi/chinesenlpAvatar de didi

    didi/ChineseNLP

    1,811Voir sur GitHub↗

    Datasets, SOTA results of every fields of Chinese NLP

    HTMLchinese-nlpchinese-word-segmentationentity-linking
    Voir sur GitHub↗1,811
  • complementizer/wcep-mds-datasetAvatar de complementizer

    complementizer/wcep-mds-dataset

    61Voir sur GitHub↗

    The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an event. These articles consist of sources cited by editors…

    Python
    Voir sur GitHub↗61
  • google-research-datasets/dakshinaAvatar de google-research-datasets

    google-research-datasets/dakshina

    212Voir sur GitHub↗

    The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages. For each language, the dataset includes a large collection of native script Wikipedia text, a romanization lexicon of words in the native script with attested romanizations, and some full sentence parallel data in both a native script of the language and the basic Latin alphabet.

    Voir sur GitHub↗212
  • kfcd/chaiziAvatar de kfcd

    kfcd/chaizi

    811Voir sur GitHub↗

    漢語拆字字典

    chinesechinese-characterscomponents
    Voir sur GitHub↗811
  • liuhuanyong/chineseembeddingAvatar de liuhuanyong

    liuhuanyong/ChineseEmbedding

    454Voir sur GitHub↗

    Chinese Embedding collection incling token ,postag ,pinyin,dependency,word embedding.中文自然语言处理向量合集,包括字向量,拼音向量,词向量,词性向量,依存关系向量.共5种类型的向量

    Python
    Voir sur GitHub↗454
  • mattdesl/dictionary-of-colour-combinationsAvatar de mattdesl

    mattdesl/dictionary-of-colour-combinations

    464Voir sur GitHub↗

    palettes from A Dictionary of Colour Combinations

    Python
    Voir sur GitHub↗464
  • oye93/chinese-nlp-corpusAvatar de OYE93

    OYE93/Chinese-NLP-Corpus

    922Voir sur GitHub↗

    Collections of Chinese NLP corpus

    Pythonchinese-nlpcorpusdatasets
    Voir sur GitHub↗922
  • panhaiqi/ancientpoetryAvatar de panhaiqi

    panhaiqi/AncientPoetry

    137Voir sur GitHub↗

    古诗词语料库

    Voir sur GitHub↗137
  • chinese-poetry/chinese-poetryAvatar de chinese-poetry

    chinese-poetry/chinese-poetry

    51,906Voir sur GitHub↗

    This project is a comprehensive dataset and archive of classical Chinese poetry, prose, and Confucian classics. It serves as a digital humanities corpus, providing machine-readable access to hundreds of thousands of poems and detailed poet biographies, specifically spanning the Tang and Song dynasties. The collection is distinguished by its scholarly depth, incorporating textual variation annotations to track disputed characters across different source editions. It also includes tonal pattern mapping to describe the rhythmic and phonetic structures of the verse, alongside a popularity ranking

    JavaScriptchinesechinese-poetryci
    Voir sur GitHub↗51,906
  • several27/fakenewscorpusAvatar de several27

    several27/FakeNewsCorpus

    413Voir sur GitHub↗

    A dataset of millions of news articles scraped from a curated list of data sources.

    artificial-intelligencecorpusdatabase
    Voir sur GitHub↗413
  • dbamman/litbankAvatar de dbamman

    dbamman/litbank

    377Voir sur GitHub↗

    Annotated dataset of 100 works of fiction to support tasks in natural language processing and the computational humanities.

    Python
    Voir sur GitHub↗377
  • facebookresearch/laserAvatar de facebookresearch

    facebookresearch/LASER

    3,659Voir sur GitHub↗

    LASER is a cross-lingual sentence embedding library and multilingual text encoder. It functions as a parallel text mining tool that maps sentences from multiple languages into a shared vector space for similarity and classification tasks. The system converts raw text into fixed-length embeddings, enabling the discovery of translation pairs by calculating the vector distance between sentences. This shared representation allows for cross-lingual document classification, where a model trained on one language can be used to categorize documents in another. The library includes a sentence-piece t

    Jupyter Notebook
    Voir sur GitHub↗3,659
  • xiangyuecn/areacity-jsspider-statsgovAvatar de xiangyuecn

    xiangyuecn/AreaCity-JsSpider-StatsGov

    6,672Voir sur GitHub↗

    This project is an administrative GIS toolset that provides a comprehensive dataset of China's administrative divisions, including provinces, cities, districts, and townships. It functions as a coordinate system transformer and a boundary converter for transforming geographic data into standard formats. The toolset distinguishes itself through the ability to convert administrative boundary data between CSV, GeoJSON, Shapefiles, and SQL. It includes specialized utilities for coordinate system transformation between GCJ-02, BD-09, WGS-84, and CGCS2000 standards to ensure accuracy across differe

    JavaScript
    Voir sur GitHub↗6,672
  • openlmlab/gaokao-benchAvatar de OpenLMLab

    OpenLMLab/GAOKAO-Bench

    760Voir sur GitHub↗

    GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.

    Python
    Voir sur GitHub↗760