16 repositorios
Toolkits for text analysis, entity extraction, and linguistic processing.
Distinguishing note: Focuses on NLP-specific dependencies for memory retrieval rather than general machine learning frameworks.
Explore 16 awesome GitHub repositories matching artificial intelligence & ml · Natural Language Processing Libraries. Refine with filters or upvote what's useful.
Mem0 is an agent-agnostic memory layer designed to provide intelligent agents with long-term persistence and cross-session state management. By acting as a centralized service, it allows diverse AI agents to recall user preferences, past interactions, and historical context, ensuring continuity across multiple workflows and independent agent systems. The platform distinguishes itself through a multi-signal retrieval engine that combines semantic vectors, keyword matching, and entity-linked metadata to surface the most relevant information. It employs an adaptive memory engine that automatical
Enables hybrid search and entity extraction through integrated natural language processing tools.
This project is a comprehensive, community-driven directory of software resources, libraries, and frameworks for the Java programming language. It serves as a centralized knowledge base designed to help developers discover tools and industry-standard solutions for building and maintaining software applications. The repository distinguishes itself through a hierarchical taxonomy that organizes a vast array of technical components into a structured, navigable tree. By relying on distributed peer contributions, the index remains a living resource that reflects current community-recommended pract
Lists Java libraries for natural language processing.
Tiktoken is a library for converting raw text into numerical sequences using byte pair encoding schemes. It functions as a toolkit for managing tokenization processes, enabling the transformation of text into the specific numerical formats required by language models. The library provides mechanisms for automated encoder selection, allowing users to retrieve the correct tokenization configuration based on specific model names. It also supports the definition and registration of custom tokenization schemes, which facilitates the use of specialized vocabularies or unique model architectures wit
Offers a toolkit for managing custom tokenization configurations and mapping text to tokens for machine learning applications.
Gensim is a natural language processing toolkit designed for large-scale text analysis and the training of semantic vector embeddings. It provides a framework for identifying latent thematic structures within document collections and calculating semantic similarity between text segments using unsupervised statistical algorithms. The project is distinguished by its ability to handle datasets that exceed available system memory through incremental corpus streaming, which processes documents one at a time from disk. It utilizes sparse vector representations and dictionary-based token mapping to
Offers a comprehensive toolkit for processing large text corpora, calculating similarity, and performing semantic analysis.
Newspaper is a Python library designed for scraping, parsing, and analyzing web-based information. It functions as a framework for automated news aggregation and large-scale web content extraction, providing tools to download, clean, and structure text, metadata, and media from diverse online sources. The project distinguishes itself through a pipeline-oriented architecture that combines heuristic-based content extraction with natural language processing. It automatically identifies and isolates article bodies from web page boilerplate while simultaneously performing language detection, keywo
Integrates natural language processing capabilities for automated keyword extraction, language detection, and text summarization of web content.
This project is a comprehensive Python toolkit designed for natural language processing, research, and education. It functions as a linguistic data processor that provides a standardized framework for managing, cleaning, and analyzing large collections of annotated text corpora and lexical resources. The library distinguishes itself through its integration of both symbolic and statistical methods, allowing users to perform complex tasks ranging from rule-based grammar parsing to machine learning-driven classification. It offers a modular pipeline for text processing, enabling the transformati
Provides a comprehensive toolkit for symbolic and statistical natural language processing, including text analysis and linguistic corpora management.
Compromise is a natural language processing library and rule-based text parser designed to analyze unstructured text. It functions as a toolkit for identifying parts of speech, linguistic patterns, and semantic meaning, while providing specialized engines for named entity recognition and the parsing of temporal and numeric data. The project is distinguished by its linguistic morphological engine, which can conjugate verbs across different tenses and inflect nouns and adjectives. It further allows for linguistic model customization through a plugin system that enables the extension of lexicons
Functions as a comprehensive toolkit for parsing unstructured text to identify parts of speech and semantic meaning.
Compromise is a natural language processing library and rule-based engine designed for English text manipulation, analysis, and parsing. It provides a toolkit for tokenizing text, identifying parts of speech, and performing linguistic analysis to achieve semantic understanding of unstructured strings. The project distinguishes itself through its ability to programmatically transform grammar, such as modifying verb tenses, noun plurality, and adjective forms. It also functions as a named entity recognizer capable of extracting people, places, organizations, dates, and contact information from
Serves as a toolkit for tokenizing text, identifying parts of speech, and performing linguistic analysis.
SentencePiece is a text segmentation engine and tokenization library designed for machine learning workflows. It provides a comprehensive toolkit for transforming raw text into subword units or numerical identifiers, enabling consistent data representation for neural network training and inference. The library supports the training of segmentation models from raw text, allowing for the creation of custom vocabularies tailored to specific domain requirements. The project distinguishes itself through its byte-level encoding and fallback mechanisms, which ensure that every input can be represent
Provides a collection of tools for normalizing, encoding, and decoding text into subword units.
This project is a conversational AI software development kit and framework used to build interactive chatbots that engage in natural language conversations and execute tasks for end users. It provides a multi-channel bot framework that connects conversational agents to various external messaging services using standardized adapters. The SDK includes a conversational workflow orchestrator and a natural language processing toolkit for analyzing user intent and extracting entities to route conversation flows. It further incorporates a speech integration framework that enables bidirectional audio
Ships a suite of NLP utilities for analyzing human language to route conversation flows and extract entities.
Smile is a comprehensive JVM machine learning library and statistical computing toolkit. It provides a suite of algorithms for classification, regression, and clustering, implemented natively for Java, Scala, and Kotlin. The project also functions as a deep learning framework, a natural language processing library, and an inference engine for large language models. The library distinguishes itself through GPU acceleration via LibTorch bindings and support for the ONNX model interchange format. It includes specialized capabilities for large language model inference, featuring Byte-Pair Encodin
Offers a full NLP library for tokenization, stemming, part-of-speech tagging, and keyword extraction.
Synonyms es una librería de procesamiento de lenguaje natural y motor de similitud semántica diseñado específicamente para texto en chino. Funciona como un kit de herramientas de word embedding y tokenizador que extrae el significado semántico e identifica sinónimos calculando la cercanía conceptual entre palabras y oraciones. El sistema proporciona un kit de herramientas para word embedding en chino y descubrimiento de sinónimos, permitiendo la recuperación de palabras semánticamente similares para expandir el vocabulario. Se distingue por un enfoque basado en configuración para la carga de modelos, que admite la integración de word embeddings personalizados para definir el espacio semántico utilizado para las búsquedas de similitud. Sus capacidades más amplias incluyen la segmentación de texto en chino con etiquetado gramatical, extracción de palabras clave y resumen de texto. La librería transforma texto sin procesar en representaciones numéricas mediante vectorización de palabras y oraciones, utilizando métricas de distancia para realizar cálculos y comparaciones de similitud semántica.
Implements a comprehensive set of NLP tools including tokenization, segmentation, and vectorization.
Este repositorio es un programa educativo integral y un framework de deep learning diseñado para enseñar aprendizaje profundo práctico usando PyTorch a través de notebooks y ejemplos de código. Sirve como una librería de alto nivel para construir, entrenar y desplegar redes neuronales, actuando como un orquestador de entrenamiento de modelos que coordina modelos de PyTorch, optimizadores y funciones de pérdida. El proyecto proporciona kits de herramientas especializados para visión artificial, procesamiento de lenguaje natural y preprocesamiento de datos tabulares. Se distingue por controles de entrenamiento avanzados como tasas de aprendizaje discriminativas, un sistema de callbacks bidireccional para personalizar la lógica de entrenamiento y una abstracción de learner de alto nivel que automatiza la colocación en dispositivos y los bucles de entrenamiento. El framework cubre una amplia superficie de capacidades, incluyendo la construcción automatizada de pipelines de datos, análisis de arquitectura de modelos y evaluación de rendimiento en tareas de clasificación, regresión y segmentación. También incluye utilidades para entrenamiento distribuido en múltiples GPUs, entrenamiento de precisión mixta para optimización de memoria y soporte especializado para datos de imágenes médicas. El proyecto se entrega como una serie de Jupyter Notebooks.
A framework for tokenizing text, managing vocabularies, and building language models and text classifiers.
OpenNRE es una librería de procesamiento de lenguaje natural y un framework de extracción de relaciones neuronales diseñado para transformar texto no estructurado en datos relacionales estructurados. Sirve como un kit de herramientas para identificar tipos de relaciones entre entidades y generar triples entidad-relación-entidad para poblar y expandir bases de conocimiento. El framework proporciona herramientas tanto para la extracción de relaciones supervisada como supervisada a distancia, permitiendo que los modelos neuronales se entrenen en datasets etiquetados o mediante pipelines automatizados que alinean triples de bases de conocimiento con texto sin procesar. El proyecto cubre un pipeline completo de extracción de información, incluyendo codificación de texto basada en transformer, inferencia de relaciones y la salida de triples estructurados para la construcción de grafos de conocimiento.
Provides a set of tools for analyzing human language to transform unstructured text into structured relational data.
This framework is a research-oriented toolkit designed for training, fine-tuning, and evaluating conversational agents using transformer-based language architectures. It provides an integrated environment for adapting large pre-trained models to specific dialogue datasets, enabling the development of systems capable of generating coherent, human-like responses. The project distinguishes itself through its support for multi-GPU distributed training, which accelerates the optimization of large-scale models. It also features configurable probabilistic decoding strategies, such as nucleus and gre
Provides a library of utilities for fine-tuning language models for interactive conversational settings.
Este proyecto es una librería estadística y un framework computacional diseñado para el modelado de temas (topic modeling) dentro de grandes colecciones de documentos. Funciona como un kit de herramientas de procesamiento de lenguaje natural que identifica estructuras temáticas ocultas analizando patrones de frecuencia de palabras a través de datos de texto no estructurados. La librería emplea Latent Dirichlet Allocation para modelar documentos como mezclas de temas y temas como mezclas de palabras. Utiliza muestreo de Gibbs y actualización iterativa del espacio de estados para estimar la distribución posterior de variables latentes, refinando las asignaciones de temas hasta que el modelo alcanza convergencia estadística. Para manejar datasets a gran escala, el framework incorpora estructuras de matrices dispersas eficientes en memoria para gestionar matrices documento-término. Proporciona capacidades para calcular y extraer distribuciones de probabilidad de temas para documentos individuales y colecciones completas, mientras incluye herramientas para monitorear el rendimiento computacional durante el análisis.
Acts as a natural language processing toolkit for analyzing and categorizing unstructured text data into thematic clusters.