awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
hankcs avatar

hankcs/HanLP

0
View on GitHub↗
36,413 स्टार्स·10,924 फोर्क्स·Python·Apache-2.0·11 व्यूज़www.hanlp.com↗

HanLP

HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion.

The project distinguishes itself through a dedicated focus on Chinese linguistic structures, including a specialized script converter for transforming text between Simplified Chinese, Traditional Chinese, and Pinyin. It further supports domain-specific model training to improve the recognition of professional terminology within specialized datasets.

Its broader capabilities cover information extraction via named entity recognition and text summarization, as well as comprehensive linguistic analysis including part-of-speech tagging and dependency syntax parsing. The toolkit also provides semantic analysis for sentiment detection and coreference resolution, alongside text transformation utilities for grammar and style conversion.

Features

  • Natural Language Processing - Provides a comprehensive toolkit for natural language processing specifically optimized for Chinese linguistic structures.
  • Text Tokenizers - Provides tools and algorithms for segmenting raw text into discrete tokens using linguistic and rule-based strategies.
  • Chinese NLP Libraries - Serves as a comprehensive natural language processing library specifically optimized for Chinese linguistic structures.
  • Deep Learning NLP Frameworks - Utilizes neural networks for advanced text analysis, sentiment detection, and semantic understanding.
  • Dependency Syntax Analyzers - Links words through directed edges to represent grammatical dependencies and structural relationships within a sentence.
  • Textual Entity Extractors - Provides automated processes for identifying and categorizing people, organizations, and locations within unstructured text.
  • Part-of-Speech Taggers - Assigns grammatical labels to words based on context and linguistic rules.
  • Script Conversion - Transforms text between Pinyin, Simplified Chinese, and Traditional Chinese characters using script conversion rules.
  • Script Converters - Transforms text between Simplified Chinese, Traditional Chinese, and Pinyin phonetic representations.
  • Semantic Analysis - Evaluates text meaning through similarity calculations, coreference resolution, and semantic role labeling.
  • Semantic Analysis Tools - Provides tools for dependency parsing, coreference resolution, and semantic role labeling to extract deep meaning.
  • Syntactic Parsers - Deconstructs sentences into hierarchical phrases and nested structures to reveal underlying grammatical organization.
  • Constituent Syntax Analysis - Deconstructs sentences into hierarchical phrases and clauses to reveal the underlying grammatical structure.
  • Deep Learning Architectures - Integrates neural network architectures to perform complex linguistic tasks such as entity recognition and syntactic parsing.
  • Dependency Syntax Analysis - Maps the grammatical relationships and dependencies between individual words within a sentence.
  • Information Extraction - Implements techniques for extracting structured data and key information from unstructured text.
  • Keyword and Phrase Extraction - Isolates the most important words and key phrases that represent the primary topic of a document.
  • Language Detection Tools - Provides utilities for identifying the specific language of provided text content.
  • Linguistic Pattern Analysis - Performs tokenization, part-of-speech tagging, and entity recognition across multiple languages to decode linguistic patterns.
  • Model Fine-Tuning - Supports adapting pre-trained models to specialized datasets to improve the recognition of professional terminology.
  • Text Classification - Groups documents into categories or clusters based on content and meaning to organize large bodies of text.
  • Multilingual Text Processing - Provides a system for tokenization, part-of-speech tagging, and named entity recognition across multiple languages.
  • Text Summarization - Provides automated methods for condensing long documents into concise summaries.
  • Semantic Similarity Calculation - Calculates the semantic relationship between two texts to determine how closely they relate in meaning.
  • Sentiment Analysis Tools - Classifies the emotional tone of text as positive, negative, or neutral.
  • Text Summarization - Condenses long documents into concise summaries and extracts key phrases to isolate important information.
  • Vector Embeddings - Represents text as high-dimensional vectors to calculate mathematical similarity between different pieces of content.
  • Coreference Resolution - Provides tools for resolving references in text to identify when different phrases refer to the same entity.
  • Domain Specific Models - Supports training deep learning models on specialized datasets to recognize professional domain terminology.
  • Custom Dictionaries - Allows defining custom word lists to force, merge, or correct how text is split into tokens.
  • Chinese NLP Toolkits - Comprehensive Java-based natural language processing library.
  • NLP Frameworks - Multilingual library for advanced natural language processing.

स्टार हिस्ट्री

hankcs/hanlp के लिए स्टार हिस्ट्री चार्टhankcs/hanlp के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

अक्सर पूछे जाने वाले प्रश्न

hankcs/hanlp क्या करता है?

HanLP is a natural language processing library and deep learning framework specifically optimized for the Chinese language, while also functioning as a multilingual text processor. It serves as a toolkit for performing linguistic analysis, semantic understanding, and script conversion.

hankcs/hanlp की मुख्य विशेषताएं क्या हैं?

hankcs/hanlp की मुख्य विशेषताएं हैं: Natural Language Processing, Text Tokenizers, Chinese NLP Libraries, Deep Learning NLP Frameworks, Dependency Syntax Analyzers, Textual Entity Extractors, Part-of-Speech Taggers, Script Conversion।

hankcs/hanlp के कुछ ओपन-सोर्स विकल्प क्या हैं?

hankcs/hanlp के ओपन-सोर्स विकल्पों में शामिल हैं: nltk/nltk — This project is a comprehensive Python toolkit designed for natural language processing, research, and education. It… isnowfy/snownlp — SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation,… ownthink/knowledgegraphdata — KnowledgeGraphData is a collection of structured datasets and corpora designed to provide a foundational layer for… flairnlp/flair — Flair is a transformer-based natural language processing framework used to build and train models for text… sloria/textblob — TextBlob is a natural language processing library that provides a unified interface for common linguistic tasks. It… fxsjy/jieba — This project is a Chinese text segmentation library and tokenizer designed to split Chinese sentences into individual…

HanLP के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो HanLP के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • nltk/nltknltk का अवतार

    nltk/nltk

    14,649GitHub पर देखें↗

    This project is a comprehensive Python toolkit designed for natural language processing, research, and education. It functions as a linguistic data processor that provides a standardized framework for managing, cleaning, and analyzing large collections of annotated text corpora and lexical resources. The library distinguishes itself through its integration of both symbolic and statistical methods, allowing users to perform complex tasks ranging from rule-based grammar parsing to machine learning-driven classification. It offers a modular pipeline for text processing, enabling the transformati

    Pythonmachine-learningnatural-language-processingnlp
    GitHub पर देखें↗14,649
  • isnowfy/snownlpisnowfy का अवतार

    isnowfy/snownlp

    6,631GitHub पर देखें↗

    SnowNLP is a Python library for Chinese natural language processing. It provides tools for text segmentation, sentiment analysis, document classification, and phonetic transliteration. The library includes capabilities for training and saving custom machine learning models for tokenization and sentiment analysis using raw training datasets. It covers a range of linguistic processing areas, including parts of speech tagging, sentence splitting, and text similarity measurement. The toolkit also provides utilities for extracting key information through text summarization and calculating word im

    Python
    GitHub पर देखें↗6,631
  • ownthink/knowledgegraphdataownthink का अवतार

    ownthink/KnowledgeGraphData

    5,181GitHub पर देखें↗

    KnowledgeGraphData is a collection of structured datasets and corpora designed to provide a foundational layer for cognitive intelligence and artificial intelligence systems. It primarily consists of large-scale Chinese knowledge graph datasets, including entity-relation data and NLP training sets used to drive semantic understanding and automated question answering. The project focuses on the construction and export of massive entity-attribute-value graphs, organizing knowledge into portable formats. It provides specialized domain partitioning to tailor information retrieval for professional

    Python
    GitHub पर देखें↗5,181
  • flairnlp/flairflairNLP का अवतार

    flairNLP/flair

    14,378GitHub पर देखें↗

    Flair is a transformer-based natural language processing framework used to build and train models for text classification and sequence tagging. It provides a specialized library for generating contextual text embeddings and performing linguistic analysis. The framework includes dedicated tools for named entity recognition, including the identification of specialized biomedical entities across multiple languages. It further supports entity linking to map identified text mentions to unique entries within general or biomedical knowledge bases. The project covers a broad range of language analys

    Python
    GitHub पर देखें↗14,378
HanLP के सभी 30 विकल्प देखें→