22 个仓库
Strategies for dividing long text into smaller overlapping segments to fit model token limits.
Distinct from Text Tokenization: Distinct from general tokenization: focuses on structural chunking with overlap for RAG context windows.
Explore 22 awesome GitHub repositories matching artificial intelligence & ml · Text Chunks. Refine with filters or upvote what's useful.
This project is a retrieval-augmented generation application designed to answer questions from uploaded PDF documents. It functions as a document question-answering engine and a streaming AI chat interface that provides responses backed by specific source citations. The system utilizes a state-machine workflow orchestrator to coordinate multi-step document ingestion and retrieval pipelines. This orchestration allows for step-by-step visualization and debugging of the process as documents are parsed and processed. The application manages the full lifecycle of document interaction, including P
Divides PDF text into smaller, overlapping segments to ensure retrieved context fits within the LLM window.
This project provides a dockerized AI workflow stack and orchestration templates for deploying a self-hosted AI environment. It establishes a localized infrastructure for building autonomous agents and model chains that process private data on-premises without external cloud dependencies. The environment is designed to support autonomous agent development, allowing models to dynamically select tools, execute shell commands, and interact with local file systems. It includes integrated vector database support to enable retrieval augmented generation and private document analysis. The stack cov
Divides long documents into smaller overlapping segments to fit model token limits for RAG pipelines.
This project is a collection of tutorials and guides for building large language model applications using the LangChain framework, written in Chinese. It serves as a learning resource for developing software that integrates language models with memory and chain-based logic. The resource provides specific walkthroughs for implementing retrieval augmented generation systems using vector stores and document loaders. It includes guides on creating autonomous agents that dynamically select and execute external tools, as well as tutorials for translating plain text queries into executable database
Provides methods for splitting documents into smaller segments to remain within token limits.
Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo
Splits text into chunks with heading paths for hierarchical context in RAG retrieval.
Verba is a retrieval-augmented generation interface and chatbot that uses Weaviate to provide factual answers based on private datasets. It functions as a vector database knowledge base, combining a hybrid search engine with an orchestration interface to connect various large language model providers and embedding services. The system differentiates itself through a RAG pipeline manager for adjusting text chunking rules and retrieval settings, alongside a 3D vector space visualization tool for analyzing the spatial organization and clustering of high-dimensional embeddings. It employs a modul
Segments large documents into smaller pieces using token or semantic rules to optimize retrieval precision.
This project is an educational implementation guide and framework for building Retrieval Augmented Generation systems. It provides a workflow for constructing a knowledge base pipeline that partitions documents, indexes them as vectors, and provides external context for language model prompts. The system features a document chunking framework that uses recursive character splitting to fit text into model context windows. It includes an in-memory vector store and a similarity search system that retrieves relevant text segments by calculating the mathematical distance between dense embedding ve
Implements strategies for dividing long text into smaller overlapping segments to fit model token limits.
pdfGPT is a retrieval augmented generation application and chatbot designed to analyze PDF documents. It functions as a document analyzer and vector search interface, using large language models to answer questions grounded in the content of uploaded files. The system implements a pipeline that extracts text from PDFs, splits content into overlapping segments, and uses vector-based semantic search to retrieve relevant context. This process allows the application to provide responses with verifiable source citations, including page number references to the original document. The project also
Divides large PDF documents into overlapping segments to optimize content for LLM context windows.
Rags is an orchestration tool for building retrieval-augmented generation pipelines and managing conversational data interfaces. It serves as a system for creating these pipelines from local files and web pages using natural language instructions to query, retrieve, and summarize information from connected datasets. The project features a multimodal retrieval system that identifies and extracts information across different data types and modalities. It includes a vector search orchestrator to manage chunking strategies and search parameters, alongside a pipeline builder that translates conver
Splits large files into overlapping segments to maintain context within model token limits for RAG.
pgai 是一个 PostgreSQL AI 工具包和框架,旨在将大语言模型和向量嵌入直接集成到数据库中。它充当了在标准数据库查询中执行机器学习模型请求和进行文本转 SQL 翻译的桥梁。 该项目提供了一个自动化的向量嵌入流水线,负责处理来自表和非结构化文档的文本加载、解析和分块。该系统利用后台工作进程在源数据发生变化时自动同步嵌入,并包含用于构建检索增强生成(RAG)应用和语义搜索引擎的专用工具。 该工具包涵盖了广泛的功能领域,包括利用 OCR 处理非结构化数据、创建将数据库模式映射到自然语言的语义目录,以及通过向量索引和结果重排序实现高性能相似度搜索。它还支持通过 SQL 调用外部模型,从而实现数据增强、分类和内容审核。
Splits long text into smaller segments using configurable algorithms and metadata injection for RAG context windows.
node-DeepResearch is an autonomous web research engine that uses large language models to iteratively search, read, and reason over web content to answer complex questions. It provides a chat-based interface that displays real-time reasoning steps and final answers, and can be configured to focus exclusively on academic papers by limiting searches to academic repositories. The research engine operates through an agentic search-read-reason loop that repeatedly searches, reads, and reasons until a stopping condition is satisfied. It enforces a token budget to cap total consumption and failed at
Selects the most relevant text segments from a document using vector similarity comparison.
Casibase is an open-source platform that orchestrates multi-turn conversations with large language models and manages retrieval-augmented knowledge bases from a single interface. It provides a unified system for connecting to over 30 AI model providers, ingesting documents into vector embeddings for semantic search, and running autonomous agent loops that can drive a browser, search the web, execute commands, and integrate with external tools. The platform distinguishes itself by combining AI conversation management with infrastructure and application orchestration capabilities. It includes a
Splits long documents into smaller segments to fit within model context windows.
Franc 是一个自然语言检测库和命令行标识符,用于确定文本样本的书写语言。它作为一个统计语言分析器,通过分析字符分布来识别和分类多语言文本。 该工具采用基于三元组(trigram)的统计分析系统,将输入样本中三个字符序列的频率与参考配置文件进行比较。它通过计算输入与这些配置文件之间的统计距离来对潜在的语言匹配进行排名,从而返回一个可能的语言排名列表。 该项目提供了一个用于文本分析的命令行界面,并支持自动化内容分类。它利用模块化语言数据集和预计算的 n-gram 表将检测逻辑与特定语言的配置文件数据分离开来。
Identifies the natural language of text samples using statistical analysis.
Chonkie 是一个专为检索增强生成 (RAG) 流水线设计的文本分块库。它充当语义文本分割器和 RAG 数据摄取流水线,将原始文本转换为嵌入片段,以便存储在向量数据库中。 该项目通过专门的分割策略脱颖而出,包括用于保留源代码逻辑边界的基于 AST 的代码分割器,以及使用嵌入模型根据语义确定边界的语义文本分割器。它还提供了一个向量数据库摄取器,用于自动化生成嵌入并将其导出到各种存储中。 该库涵盖了广泛的功能,包括通过 OCR 和 Markdown 提取进行文档解析,多种分割方法(如基于 Token 计数和分层分割),以及通过可重用流水线进行工作流编排。它支持多种向量存储集成,包括 Qdrant、Milvus、Weaviate 和 Elasticsearch,以及将数据导出为 JSON 和 Hugging Face 数据集。 用户可以通过命令行界面执行这些操作,或将系统部署为容器化的 API 服务。
Provides a comprehensive library for dividing documents into semantic, structural, or token-based chunks for RAG pipelines.
AdalFlow 是一个自主 AI 代理框架和 LLM 应用库,旨在构建模块化工作流。它作为一个模型无关的接口和 RAG 流水线编排器,允许用户开发 ReAct 代理,利用迭代推理和外部工具执行来解决复杂任务。 该项目通过一个提示词优化系统脱颖而出,该系统使用文本梯度下降自动优化提示词模板和少样本示例。它将模型反馈视为可微分信号,实现了一种 LLM 反向传播形式,从而根据评估指标迭代提高输出质量。 该框架涵盖了广泛的功能面,包括带有语义向量搜索和重排序的检索增强生成、用于可观测性的基于跨度的执行追踪,以及模式驱动的结构化解析。它为众多专有和开源模型提供商提供了统一的通信层,并支持将 Python 函数转换为标准化的工具接口。 该系统使用 Python 实现,并与 MLflow 集成以进行工作流跟踪和分析。
Divides large text into smaller, overlapping segments using tokenizers to fit within model context windows.
Langroid is a multi-agent orchestration framework and tool integration suite designed for building complex AI applications. It serves as a multi-modal integration layer that connects diverse local and remote language models with an agentic retrieval-augmented generation system. The project distinguishes itself through a collaborative message-exchange paradigm, allowing specialized agents to delegate tasks hierarchically and coordinate via structured communication. It features an advanced state management system for conversational AI, including the ability to rewind and prune conversation hist
Divides long text into segments of specific token lengths while respecting linguistic boundaries.
nano-graphrag 是一个检索系统,使用知识图谱为大语言模型响应提供结构化上下文。它既是一个将非结构化文本转换为实体和关系网络的知识图谱索引器,也是一个混合图检索系统。 该项目通过结合局部邻域搜索和全局社区摘要来回答复杂的自然语言问题,从而脱颖而出。它包含一个知识图谱可视化工具,可生成实体及其关系的 HTML 表示,以映射索引知识。 该框架涵盖了广泛的功能,包括实体关系提取、基于社区的图聚类和基于哈希的增量索引。它提供了一个集成层,用于连接开源模型和本地嵌入提供程序,并支持用于键值、向量和图数据的可插拔存储后端。通过基于参数的响应缓存和用于修复语言模型不稳定 JSON 输出的后处理函数,提供了额外的实用性。
Provides configurable logic to divide raw text into segments based on token limits or delimiters for RAG.
Tika is a content analysis toolkit and Java library designed for detecting and extracting metadata and text from thousands of different file types. It functions as a universal document text extractor and metadata extraction engine, converting complex files into plain text or XHTML. The system employs a specialized MIME type detector that identifies document formats using magic bytes and metadata to determine the correct parser. It serves as an OCR integration gateway, connecting to external text recognition tools to extract content from image files. The project covers a broad range of extrac
Processes extracted content in small segments via custom handlers to maintain a low memory footprint.
mcp-context-forge is a Model Context Protocol federation gateway that unifies diverse AI tool servers and APIs into a single consistent interface for discovery and execution. It acts as a centralized proxy that aggregates multiple servers and APIs, allowing AI agents to access and invoke a unified set of tools, prompts, and resources. The project distinguishes itself through a multi-protocol translation bridge that converts communication between standard I/O, SSE, gRPC, and REST to enable interoperability between disparate tool servers. It includes a comprehensive LLM evaluation framework for
Evaluates text statistics to recommend effective semantic chunking strategies for LLM context windows.
pg_textsearch is a full-text search integration for PostgreSQL that provides large-scale text indexing and BM25 relevance ranking. It implements a scalable indexing architecture that uses a memtable system to spill data to disk segments, allowing for the processing of massive datasets. The project distinguishes itself through support for multilingual search via language-specific partial indexes and the ability to index complex expressions, such as JSONB fields or concatenated columns. It ensures high availability by utilizing PostgreSQL-native streaming replication and write-ahead logs to syn
Splits oversized text bodies into smaller pieces during tokenization to maintain memory efficiency and consistency.
This project is a document digitization utility that combines traditional optical character recognition with language model processing to convert scanned PDF files into structured markdown. It functions as an automated pipeline that extracts raw text from images and applies intelligent post-processing to refine the output. The system distinguishes itself by using language models to perform error correction, removing artifacts and formatting inconsistencies common in raw character recognition. It incorporates a modular design that decouples processing logic from specific model providers, allow
Splits long documents into overlapping segments to ensure language models maintain semantic coherence while respecting token limits.