36 रिपॉजिटरी
Systems for identifying and classifying entities such as people, organizations, and locations within unstructured text.
Distinct from Named Entity Recognition: Existing candidates were restricted to Awesome Lists or specific narrow extraction tasks.
Explore 36 awesome GitHub repositories matching artificial intelligence & ml · Named Entity Recognition. Refine with filters or upvote what's useful.
Compromise is a natural language processing library and rule-based text parser designed to analyze unstructured text. It functions as a toolkit for identifying parts of speech, linguistic patterns, and semantic meaning, while providing specialized engines for named entity recognition and the parsing of temporal and numeric data. The project is distinguished by its linguistic morphological engine, which can conjugate verbs across different tenses and inflect nouns and adjectives. It further allows for linguistic model customization through a plugin system that enables the extension of lexicons
Identifies and extracts specific categories of information like people, organizations, and dates from unstructured text.
CoreNLP is a Java natural language processing library designed to convert raw human language text into structured data. It utilizes a suite of linguistic annotators to analyze text through a pipeline, extracting grammatical structures, sentiment, and linguistic patterns. The project includes a coreference resolution engine that links multiple mentions of the same entity to maintain contextual consistency across documents. It also provides tools for named entity recognition to categorize people, companies, and locations, and a part-of-speech tagger to assign grammatical categories and base for
Identifies and categorizes named entities using pattern matching and gazetteers.
AutoGluon is an automated machine learning framework and multimodal library designed to automate the end-to-end pipeline from data preprocessing to high-accuracy model training and validation. It functions as an automated model trainer for tabular, image, text, and time series data, as well as a tool for time series forecasting and foundation model finetuning. The project is distinguished by its ability to jointly process and fuse different data types, allowing for the construction of multimodal neural networks that integrate images, text, and structured tables. It supports zero-shot inferenc
Identifies and classifies key entities such as people and organizations within text strings across multiple languages.
ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It
Classifies entities from transcripts by combining speech recognition with natural language processing.
Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo
Identifies people, organizations, locations, and other entities in extracted text using ONNX or LLM providers.
This project is a comprehensive collection of educational examples and reference implementations for building vision and language models using PyTorch. It serves as a deep learning tutorial covering the end-to-end process of developing neural networks, from initial architecture definition to final production deployment. The repository provides detailed guides on implementing a wide range of domain-specific models, including convolutional neural networks for object detection and segmentation, as well as transformer and recurrent architectures for natural language processing. It emphasizes gene
Implements transformer-based models to identify and classify key entities within unstructured text.
Stanza is a Python natural language processing library designed for tokenization, lemmatization, and dependency parsing across many human languages using neural models. It provides a neural processing pipeline that converts raw text into structured linguistic data objects, alongside a specialized analyzer for extracting medical insights from clinical and biomedical language. The project includes a wrapper that connects Python scripts to Java-based natural language processing tools and remote annotation servers. This enables a bridge for extracting linguistic annotations and analysis data from
Identifies and classifies entities like people, organizations, and locations within raw text.
DeepPavlov is a conversational AI framework and deep learning NLP library designed for building end-to-end dialogue systems and chatbots. It functions as an NLP pipeline orchestrator that allows users to compose pre-trained models and text processing components into sequential data flows for complex linguistic tasks. The system is distinguished by its ability to act as a chatbot deployment server, exposing trained conversational models as web services via REST and Socket APIs. It utilizes JSON-based pipeline configurations and dynamic variable interpolation to decouple model logic from infras
Identifies and categorizes specific entities and their relationships within text for information extraction.
nlp.js is a JavaScript natural language processing library and development framework used to build natural language understanding engines. It provides a toolkit for creating local machine learning models for intent classification and acts as a multilingual text processor that detects languages and normalizes text across various dialects. The framework distinguishes itself by supporting local execution on both servers and mobile devices, enabling chatbot functionality without an internet connection. It features a specialized system for conversational slot filling to collect mandatory informati
Identifies and extracts structured data like dates, currency, and contact information from unstructured text.
ansj_seg is a Java NLP toolkit and segmentation library designed for processing Chinese text. It functions as a word segmenter, part-of-speech tagger, and named entity recognizer to divide continuous Chinese characters into meaningful words and tokens. The library utilizes statistical models for text segmentation and provides capabilities for identifying and extracting person names from unstructured documents. It also assigns grammatical categories to tokens to determine their linguistic roles within a sentence. The toolkit supports domain-specific text processing through the use of custom d
Identifies and classifies entities such as person names within unstructured Chinese text.
nlp-recipes is a collection of implementation guides and reference templates for applying natural language processing techniques to real-world tasks. It provides standardized workflows and code examples for developing NLP pipelines, from dataset preparation and model training to performance evaluation. The project focuses on the practical application of transformer-based models, offering patterns for fine-tuning pretrained architectures for tasks such as text classification, named entity recognition, and question answering. It also includes a toolkit for model interpretability, allowing users
Identifies and categorizes key entities like people or locations within unstructured text.
This is a Chinese natural language processing toolkit providing a suite of tools for word segmentation, part-of-speech tagging, and named entity recognition. It includes a neural dependency parser for analyzing syntactic and semantic relationships between words and a machine learning training suite for creating custom linguistic models using annotated datasets. The toolkit distinguishes itself through its deployment flexibility, offering a dockerized server and a web service interface that exposes processing capabilities via API. It supports the use of pretrained models and allows for the int
Identifies and categorizes specific entities such as people, places, and organizations within Chinese text.
Knwl.js is a JavaScript named entity recognition library and rule-based text parser. It serves as an extensible information extraction tool designed to identify and pull structured entities, such as dates, times, and locations, from unstructured text strings. The library allows for the definition of specialized rules and custom plugins to identify and extract unique pieces of information. This extensibility enables the automation of information retrieval by converting human-readable text into structured formats for applications and databases. The system utilizes regular expression matching a
Identifies and classifies entities such as dates, times, and locations within unstructured text strings.
Knwl.js is a JavaScript named entity recognition library and text entity extractor. It functions as an extensible text parsing engine designed to scan unstructured strings for specific data patterns and convert them into structured information. The engine utilizes a modular framework that allows for the recognition of custom data types. This is achieved through a plugin system for language pattern matching, enabling the integration of custom logic to identify unique data types within text. The library identifies and isolates entities such as dates, times, phone numbers, emails, and locations
Provides a JavaScript library for identifying and classifying entities like dates, times, and locations.
KnowledgeGraphData is a collection of structured datasets and corpora designed to provide a foundational layer for cognitive intelligence and artificial intelligence systems. It primarily consists of large-scale Chinese knowledge graph datasets, including entity-relation data and NLP training sets used to drive semantic understanding and automated question answering. The project focuses on the construction and export of massive entity-attribute-value graphs, organizing knowledge into portable formats. It provides specialized domain partitioning to tailor information retrieval for professional
Identifies and classifies named entities such as people and organizations within unstructured text.
यह प्रोजेक्ट एक नेचुरल लैंग्वेज प्रोसेसिंग सिस्टम है जिसे नेम्ड एंटिटी रिकग्निशन और टेक्स्ट क्लासिफिकेशन के लिए डिज़ाइन किया गया है। यह कच्चे टेक्स्ट से विशिष्ट नामों और प्रमुख जानकारी की पहचान करने के लिए मशीन लर्निंग दृष्टिकोण का उपयोग करता है। यह सिस्टम एक मल्टी-लेयर आर्किटेक्चर लागू करता है जो एम्बेडिंग के लिए प्री-ट्रेन्ड ट्रांसफॉर्मर, सीक्वेंस मॉडलिंग के लिए बाईडायरेक्शनल लॉन्ग शॉर्ट-टर्म मेमोरी और लेबल ट्रांज़िशन के लिए कंडीशनल रैंडम फील्ड को जोड़ता है। यह प्रोजेक्ट विशिष्ट डेटासेट पर इन मॉडल्स को फाइन-ट्यून करके ट्रांसफर लर्निंग का समर्थन करता है। इसमें कस्टम डेटासेट पर मॉडल्स को प्रशिक्षित करने और प्रशिक्षित मॉडल को नेटवर्क सर्विस के रूप में डिप्लॉय करने की क्षमताएं शामिल हैं।
Identifies and classifies specific names and key information from unstructured text.
यह प्रोजेक्ट एक नेम्ड एंटिटी रिकग्निशन फ्रेमवर्क और TensorFlow-आधारित नेचुरल लैंग्वेज प्रोसेसिंग मॉडल है। यह प्री-ट्रेन्ड भाषा मॉडल्स को विशिष्ट एंटिटी रिकग्निशन और टेक्स्ट क्लासिफिकेशन कार्यों के अनुकूल बनाने के लिए एक पाइपलाइन प्रदान करता है। यह सिस्टम एक सीक्वेंस लेबलिंग आर्किटेक्चर लागू करता है जो ट्रांसफॉर्मर-आधारित एम्बेडिंग को बाईडायरेक्शनल सीक्वेंस मॉडलिंग और कंडीशनल रैंडम फील्ड डिकोडिंग के साथ जोड़ता है। इसमें मॉडल वेट्स को फाइन-ट्यून करने और असंरचित टेक्स्ट के भीतर एंटिटीज़ की पहचान और वर्गीकरण करने के लिए नेटवर्क को प्रशिक्षित करने के लिए टूल्स शामिल हैं। फ्रेमवर्क में एक क्लाइंट-सर्वर आर्किटेक्चर भी शामिल है जो प्रशिक्षित मॉडल्स को HTTP API के माध्यम से उजागर करता है।
Identifies and categorizes specific entities within unstructured text using BERT and BiLSTM-CRF models.
This project is a collection of transformer natural language processing tutorial notebooks and educational resources. It provides a guide for using the Hugging Face Transformers library through interactive coding exercises and demonstrations. The repository contains ready-to-run Jupyter notebooks that provide practical examples for implementing transformer models. These resources demonstrate how to execute specific natural language processing workflows using pre-trained models. The notebooks cover a range of natural language processing tasks, including text classification, automatic text sum
Provides examples of named entity recognition to extract key entities from unstructured text.
SwiftOCR एक नेटिव Swift लाइब्रेरी है जिसे छवियों से टेक्स्ट और अल्फ़ान्यूमेरिक अक्षरों को निकालने के लिए डिज़ाइन किया गया है। यह एक न्यूरल नेटवर्क टेक्स्ट रिकॉग्नाइज़र के रूप में कार्य करता है जो विज़ुअल डेटा से अक्षरों और स्ट्रिंग्स की पहचान करता है। इस लाइब्रेरी में एक कस्टम OCR मॉडल ट्रेनर और कस्टम फॉन्ट रिकग्निशन के लिए टूल शामिल हैं। ये क्षमताएं मान्यता सटीकता में सुधार करने के लिए विशिष्ट फॉन्ट और कैरेक्टर सेट के अनुरूप विशेष न्यूरल नेटवर्क के निर्माण की अनुमति देती हैं। यह सिस्टम व्यक्तिगत अक्षर क्षेत्रों की पहचान करने के लिए कनेक्टेड-कंपोनेंट लेबलिंग का उपयोग करता है और छोटे अल्फ़ान्यूमेरिक कोड को डिजिटल टेक्स्ट में बदलने के लिए इमेज प्रोसेसिंग का उपयोग करता है।
Enables generation of custom neural networks for new fonts and character sets.
This project is a comprehensive instructional resource and course for building neural networks using PyTorch. It covers the fundamental building blocks of deep learning, including tensor manipulation, automatic differentiation, and the construction of modular neural network components. The repository serves as a technical guide for several specialized domains. It provides implementation details for computer vision tasks such as image classification, object detection, and semantic segmentation, as well as natural language processing workflows involving transformers, recurrent networks, and gen
Implements systems for identifying and classifying entities such as people and locations within unstructured text.