awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
hankcs avatar

hankcs/HanLP

0
View on GitHub↗
36,413 estrellas·10,924 forks·Python·Apache-2.0·11 vistaswww.hanlp.com↗

HanLP

HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion.

The project distinguishes itself through a dedicated focus on Chinese linguistic structures, including a specialized script converter for transforming text between Simplified Chinese, Traditional Chinese, and Pinyin. It further supports domain-specific model training to improve the recognition of professional terminology within specialized datasets.

Its broader capabilities cover information extraction via named entity recognition and text summarization, as well as comprehensive linguistic analysis including part-of-speech tagging and dependency syntax parsing. The toolkit also provides semantic analysis for sentiment detection and coreference resolution, alongside text transformation utilities for grammar and style conversion.

Features

  • Natural Language Processing - Provides a comprehensive toolkit for natural language processing specifically optimized for Chinese linguistic structures.
  • Text Tokenizers - Provides tools and algorithms for segmenting raw text into discrete tokens using linguistic and rule-based strategies.
  • Chinese NLP Libraries - Serves as a comprehensive natural language processing library specifically optimized for Chinese linguistic structures.
  • Deep Learning NLP Frameworks - Utilizes neural networks for advanced text analysis, sentiment detection, and semantic understanding.
  • Dependency Syntax Analyzers - Links words through directed edges to represent grammatical dependencies and structural relationships within a sentence.
  • Textual Entity Extractors - Provides automated processes for identifying and categorizing people, organizations, and locations within unstructured text.
  • Part-of-Speech Taggers - Assigns grammatical labels to words based on context and linguistic rules.
  • Script Conversion - Transforms text between Pinyin, Simplified Chinese, and Traditional Chinese characters using script conversion rules.
  • Script Converters - Transforms text between Simplified Chinese, Traditional Chinese, and Pinyin phonetic representations.
  • Semantic Analysis - Evaluates text meaning through similarity calculations, coreference resolution, and semantic role labeling.
  • Semantic Analysis Tools - Provides tools for dependency parsing, coreference resolution, and semantic role labeling to extract deep meaning.
  • Syntactic Parsers - Deconstructs sentences into hierarchical phrases and nested structures to reveal underlying grammatical organization.
  • Constituent Syntax Analysis - Deconstructs sentences into hierarchical phrases and clauses to reveal the underlying grammatical structure.
  • Deep Learning Architectures - Integrates neural network architectures to perform complex linguistic tasks such as entity recognition and syntactic parsing.
  • Dependency Syntax Analysis - Maps the grammatical relationships and dependencies between individual words within a sentence.
  • Information Extraction - Implements techniques for extracting structured data and key information from unstructured text.
  • Keyword and Phrase Extraction - Isolates the most important words and key phrases that represent the primary topic of a document.
  • Language Detection Tools - Provides utilities for identifying the specific language of provided text content.
  • Linguistic Pattern Analysis - Performs tokenization, part-of-speech tagging, and entity recognition across multiple languages to decode linguistic patterns.
  • Model Fine-Tuning - Supports adapting pre-trained models to specialized datasets to improve the recognition of professional terminology.
  • Text Classification - Groups documents into categories or clusters based on content and meaning to organize large bodies of text.
  • Multilingual Text Processing - Provides a system for tokenization, part-of-speech tagging, and named entity recognition across multiple languages.
  • Text Summarization - Provides automated methods for condensing long documents into concise summaries.
  • Semantic Similarity Calculation - Calculates the semantic relationship between two texts to determine how closely they relate in meaning.
  • Sentiment Analysis Tools - Classifies the emotional tone of text as positive, negative, or neutral.
  • Text Summarization - Condenses long documents into concise summaries and extracts key phrases to isolate important information.
  • Vector Embeddings - Represents text as high-dimensional vectors to calculate mathematical similarity between different pieces of content.
  • Coreference Resolution - Provides tools for resolving references in text to identify when different phrases refer to the same entity.
  • Domain Specific Models - Supports training deep learning models on specialized datasets to recognize professional domain terminology.
  • Custom Dictionaries - Allows defining custom word lists to force, merge, or correct how text is split into tokens.
  • Chinese NLP Toolkits - Comprehensive Java-based natural language processing library.
  • NLP Frameworks - Multilingual library for advanced natural language processing.

Historial de estrellas

Gráfico del historial de estrellas de hankcs/hanlpGráfico del historial de estrellas de hankcs/hanlp

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a HanLP

Proyectos open-source similares, clasificados según cuántas características comparten con HanLP.
  • nltk/nltkAvatar de nltk

    nltk/nltk

    14,649Ver en GitHub↗

    This project is a comprehensive Python toolkit designed for natural language processing, research, and education. It functions as a linguistic data processor that provides a standardized framework for managing, cleaning, and analyzing large collections of annotated text corpora and lexical resources. The library distinguishes itself through its integration of both symbolic and statistical methods, allowing users to perform complex tasks ranging from rule-based grammar parsing to machine learning-driven classification. It offers a modular pipeline for text processing, enabling the transformati

    Pythonmachine-learningnatural-language-processingnlp
    Ver en GitHub↗14,649
  • isnowfy/snownlpAvatar de isnowfy

    isnowfy/snownlp

    6,631Ver en GitHub↗

    SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation, sentiment analysis, document classification, and phonetic transliteration. The library includes capabilities for training and saving custom machine learning models for tokenization and sentiment analysis using raw training datasets. It covers a range of linguistic processing areas, including parts of speech tagging, sentence splitting, and text similarity measurement. The toolkit also provides utilities for extracting key information through text summarization and calculating word im

    Python
    Ver en GitHub↗6,631
  • ownthink/knowledgegraphdataAvatar de ownthink

    ownthink/KnowledgeGraphData

    5,181Ver en GitHub↗

    KnowledgeGraphData is a collection of structured datasets and corpora designed to provide a foundational layer for cognitive intelligence and artificial intelligence systems. It primarily consists of large-scale Chinese knowledge graph datasets, including entity-relation data and NLP training sets used to drive semantic understanding and automated question answering. The project focuses on the construction and export of massive entity-attribute-value graphs, organizing knowledge into portable formats. It provides specialized domain partitioning to tailor information retrieval for professional

    Python
    Ver en GitHub↗5,181
  • flairnlp/flairAvatar de flairNLP

    flairNLP/flair

    14,378Ver en GitHub↗

    Flair is a transformer-based natural language processing framework used to build and train models for text classification and sequence tagging. It provides a specialized library for generating contextual text embeddings and performing linguistic analysis. The framework includes dedicated tools for named entity recognition, including the identification of specialized biomedical entities across multiple languages. It further supports entity linking to map identified text mentions to unique entries within general or biomedical knowledge bases. The project covers a broad range of language analys

    Python
    Ver en GitHub↗14,378
Ver las 30 alternativas a HanLP→

Preguntas frecuentes

¿Qué hace hankcs/hanlp?

HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion.

¿Cuáles son las características principales de hankcs/hanlp?

Las características principales de hankcs/hanlp son: Natural Language Processing, Text Tokenizers, Chinese NLP Libraries, Deep Learning NLP Frameworks, Dependency Syntax Analyzers, Textual Entity Extractors, Part-of-Speech Taggers, Script Conversion.

¿Qué alternativas de código abierto existen para hankcs/hanlp?

Las alternativas de código abierto para hankcs/hanlp incluyen: nltk/nltk — This project is a comprehensive Python toolkit designed for natural language processing, research, and education. It… isnowfy/snownlp — SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation,… ownthink/knowledgegraphdata — KnowledgeGraphData is a collection of structured datasets and corpora designed to provide a foundational layer for… flairnlp/flair — Flair is a transformer-based natural language processing framework used to build and train models for text… sloria/textblob — TextBlob is a natural language processing library that provides a unified interface for common linguistic tasks. It… fxsjy/jieba — This project is a Chinese text segmentation library and tokenizer designed to split Chinese sentences into individual…