22 repositorios
Fetching and converting web content into a format optimized for efficient use by language models.
Distinct from CMS Content Fetching: Distinct from CMS Content Fetching: focuses on fetching web content for LLM consumption, not specifically from content management systems.
Explore 22 awesome GitHub repositories matching data & databases · Web Content Fetching. Refine with filters or upvote what's useful.
Claude Code is a command-line interface and multi-agent orchestration framework designed for autonomous software engineering. It enables AI agents to perform codebase modifications, debugging, and Git workflow management while coordinating multiple specialized agents to decompose and execute complex engineering tasks in parallel. The system distinguishes itself through a high degree of isolation and safety, utilizing Git worktrees to create independent working directories for concurrent agents and implementing a tiered permission system that combines user rules, project policies, and OS-level
Downloads web content from URLs and converts it to Markdown for LLM processing.
gpt-oss is an open-weight large language model and reasoning engine designed for complex reasoning and agentic workflows. It functions as an AI agent framework and model serving API, allowing for local deployment and the hosting of standardized interfaces to expose model completions and internal reasoning processes. The project distinguishes itself as a quantized inference engine, utilizing tensor parallelism and weight quantization to run high-parameter models on limited hardware. It features a reasoning model that employs chain-of-thought processing to solve multi-step logical tasks. The s
Fetches and converts web page content into a format optimized for LLM consumption and citation.
Deep Searcher is an open-source retrieval-augmented generation engine that indexes private documents into a vector database and uses large language models to answer complex questions with cited reasoning. It functions as both a command-line interface and a web API research tool, enabling users to load data and generate comprehensive reports by combining indexed private information with LLM-powered analysis. The system distinguishes itself through a plugin-based provider architecture that supports multiple embedding models, LLM providers, vector databases, and file loaders as interchangeable c
Fetches and indexes content from specified URLs using configurable web crawlers for inclusion in the knowledge base.
Yao is an LLM agent framework and low-code web app builder designed for orchestrating autonomous AI agents. It provides a platform to design, deploy, and coordinate agents with specialized personas that can plan tasks, utilize external tools, and execute multi-stage pipelines. The project distinguishes itself through a Model Context Protocol server for connecting assistants to external binaries and HTTP services, and a gRPC remote execution engine that allows agents to manage remote servers and devices. It includes a model-agnostic provider bridge that supports dynamic switching between vario
Retrieves data, prices, and pages from external websites using a built-in scraping service optimized for LLM use.
Horizon es un sistema de agregación de noticias impulsado por IA diseñado para construir tuberías personalizadas que obtienen, filtran y enriquecen información de diversas fuentes web. Utiliza modelos de lenguaje de gran tamaño para automatizar el filtrado de información, puntuando el contenido para eliminar el ruido y resaltar historias de alto valor. El sistema integra el Protocolo de Contexto de Modelo (Model Context Protocol) para exponer las etapas de la tubería como herramientas para asistentes de IA externos. Emplea un adaptador unificado para estandarizar diversos proveedores de modelos de IA para tareas consistentes de puntuación y resumen de contenido. La tubería agrega datos de feeds RSS, plataformas sociales, kits de herramientas financieras y repositorios de código. Gestiona el contenido mediante deduplicación, filtrado de categorías basado en cuotas y enriquecimiento contextual antes de entregar resúmenes multilingües por correo electrónico, webhooks o despliegue de sitio estático. Los flujos de trabajo se orquestan a través de automatización en la nube recurrente para gestionar la recolección y entrega programada de información procesada.
Builds custom workflows to fetch, deduplicate, and enrich data from diverse web sources before final delivery.
Airweave is a unified AI knowledge base platform that syncs data from external APIs into a searchable layer for retrieval-augmented generation. It provides a pre-built data connector library and a framework for building custom connectors, enabling the extraction, transformation, and synchronization of structured and unstructured data from SaaS applications. The platform includes a hybrid vector retrieval system that combines semantic, neural, and keyword search strategies to deliver grounded context for AI agents. The platform distinguishes itself through an agentic search engine that iterati
Fetches web content from ClinicalTrials.gov pages for downstream LLM consumption.
Rainmeter is a Windows desktop widget engine that renders customizable skins and interactive widgets directly on the desktop, supporting live data feeds and user interaction. It functions as a desktop customization platform and skin authoring framework, allowing users to create widgets by defining data sources and visual elements with full style and layout control. The engine includes a Lua scripting runtime for extending widget functionality with custom logic and data processing, and provides a plugin SDK with a C/C++ API for building native plugins that add new data sources or rendering capa
Fetches and parses web content from URLs and RSS feeds for display in desktop widgets.
Fetches web content and converts it for use by language models in the IDA Pro context.
node-DeepResearch is an autonomous web research engine that uses large language models to iteratively search, read, and reason over web content to answer complex questions. It provides a chat-based interface that displays real-time reasoning steps and final answers, and can be configured to focus exclusively on academic papers by limiting searches to academic repositories. The research engine operates through an agentic search-read-reason loop that repeatedly searches, reads, and reasons until a stopping condition is satisfied. It enforces a token budget to cap total consumption and failed at
Fetches web pages, extracts clean markdown text, and generates image captions for language model ingestion.
TAICHI-flet es un navegador de recursos integrado con IA y una aplicación de escritorio para Windows construida con Flet. Sirve como un centro multimedia centralizado y agregador de contenido web diseñado para combinar utilidades de inteligencia artificial con herramientas para buscar y acceder a películas, música y software. La aplicación permite la agregación de recursos de múltiples fuentes, incluyendo unidades de almacenamiento en la nube y direcciones web externas. Proporciona herramientas especializadas para transmitir y descargar anime y música, leer novelas en línea con reproducción de texto a voz y automatizar operaciones en el sistema operativo Windows utilizando inteligencia artificial. La interfaz incluye un sistema de navegación basado en pestañas para cambiar entre categorías de contenido y un sistema de gestión de temas para personalizar la estética del escritorio y los fondos de pantalla. Las capacidades técnicas incluyen el uso de servidores proxy para saltar restricciones de seguridad de origen cruzado para imágenes remotas y procesamiento mediante hilos demonio para mantener la capacidad de respuesta de la interfaz durante tareas de larga duración.
Downloads HTML and binary data from web addresses to retrieve external multimedia resources.
Crawler4j is a multi-threaded Java web crawler and spider designed for high-volume web traversal and content extraction. It functions as a polite crawling framework that enables the discovery and indexing of HTML and binary content across multiple websites. The project distinguishes itself through a persistent crawling model that serializes session state to local storage, allowing the engine to resume indexing after a crash or interruption. It includes a politeness controller to regulate request frequency and delays, preventing server overloading and IP blocking. The system covers a broad ra
Provides control over whether to follow redirects, include encrypted pages, or process specific content types.
iflow-cli is a command-line interface and suite of AI tools designed for software engineering, workflow orchestration, and multimodal data analysis. It functions as an LLM command line interface that enables users to execute AI workflows, analyze codebase structures, and interact with large language models directly from the terminal. The project features a plugin-based agent architecture that allows for the integration of specialized domain experts and custom instruction sets from an external marketplace. It distinguishes itself through a multimodal AI terminal capable of processing visual da
Extracts raw content from specific URLs to provide text optimized for analysis by language models.
Templater is an Obsidian template engine and JavaScript automation plugin that functions as a dynamic content generator and workflow orchestrator. It enables the automation of document creation and note-taking tasks through the use of dynamic placeholders and embedded logic. The project distinguishes itself by executing custom JavaScript and shell commands to manipulate files and insert data. It allows for interactive note generation via modal prompts for user input and the import of external JavaScript modules to provide reusable logic outside of template files. Its capabilities include pro
Executes HTTP requests to retrieve remote web content for use within documents.
ollama-js es una librería cliente de JavaScript y wrapper de API que proporciona una interfaz programática para interactuar y gestionar modelos de lenguaje grandes. Permite la ejecución de modelos tanto en entornos locales como basados en la nube, facilitando la generación de texto conversacional y la gestión de los ciclos de vida de los modelos. El proyecto se distingue por ofrecer herramientas especializadas para la administración de modelos, incluida la capacidad de descargar, crear y eliminar modelos, así como la capacidad de definir blueprints de modelos personalizados y plantillas de prompts. También proporciona un cliente de incrustación vectorial (vector embedding) para generar representaciones de texto numéricas para admitir pipelines de búsqueda y recuperación semántica. La librería cubre una amplia gama de capacidades, incluyendo análisis multimodal para procesar imágenes, la captura de trazas de razonamiento interno y la aplicación de esquemas JSON estructurados para la extracción de datos. Además, admite la interacción avanzada con modelos a través de la invocación de herramientas y el streaming de respuestas a través de generadores asíncronos. La librería está escrita en TypeScript.
Retrieves raw content from specified URLs for use within language model workflows.
OpenSquilla es un framework de orquestación de agentes LLM diseñado para coordinar flujos de trabajo de IA de varios pasos y la ejecución de herramientas mediante grafos acíclicos dirigidos. Funciona como un sistema centralizado para gestionar paquetes de habilidades especializadas y ejecutar secuencias de razonamiento complejas. El proyecto se distingue por una pasarela de enrutamiento que dirige las tareas a diferentes proveedores de IA según la complejidad, el coste y el rendimiento. Utiliza un sistema de memoria de IA de varios niveles que organiza el conocimiento de trabajo, episódico y semántico mediante embeddings locales y SQLite, junto con un sandbox de ejecución seguro que aísla el código generado por el agente mediante perfiles de permisos basados en riesgos. La plataforma cubre una amplia gama de capacidades, incluyendo despliegue multicanal en web y plataformas de mensajería, programación automatizada de tareas mediante cron y un puente de Model Context Protocol para conectar con herramientas externas. También proporciona herramientas integrales de monitoreo y observabilidad para rastrear costes de tokens, auditar decisiones en tiempo de ejecución y gestionar un catálogo de habilidades reutilizables. El sistema incluye utilidades de línea de comandos para la inicialización del espacio de trabajo y la gestión del ciclo de vida de las habilidades.
Deno AI Agent reads the full content of a specific URL to perform deep inspection of a page.
OptiLLM es un proxy de inferencia y router de puerta de enlace que dirige los prompts a modelos de lenguaje específicos basados en costo, rendimiento y salud del proveedor. Funciona como una capa de middleware diseñada para optimizar las solicitudes mediante enrutamiento inteligente, balanceo de carga y gestión de contexto. El proyecto proporciona capacidades especializadas para la protección de datos mediante la anonimización de información de identificación personal antes de que las solicitudes lleguen a un modelo. También actúa como un orquestador de razonamiento y capa de integración de herramientas, utilizando bucles de tiempo de inferencia y autorreflexión para mejorar la precisión mientras conecta los modelos a servidores de protocolos externos, contenido web e intérpretes de código. La funcionalidad adicional incluye una interfaz basada en esquemas para generar salidas estructuradas legibles por máquina. El sistema también gestiona la alta disponibilidad mediante balanceo de carga a nivel de proveedor y monitoreo de salud.
Fetches and converts web content from specified URLs into a format optimized for LLM context injection.
This project is a web-based manga and novel downloader and multi-site web scraper designed to extract images and text from diverse media platforms. It functions as a digital media archiver and EPUB e-book generator, using a plugin-based crawler architecture with site-specific scripts to define how content is extracted from various international websites. The system distinguishes itself through authenticated web crawling, using browser cookie simulation to access restricted or member-only content. It includes specialized capabilities for digital comic archiving, which organizes image sequences
Implements a sequential workflow that fetches raw web content and processes it into structured archives and e-books.
Langroid is a multi-agent orchestration framework and tool integration suite designed for building complex AI applications. It serves as a multi-modal integration layer that connects diverse local and remote language models with an agentic retrieval-augmented generation system. The project distinguishes itself through a collaborative message-exchange paradigm, allowing specialized agents to delegate tasks hierarchically and coordinate via structured communication. It features an advanced state management system for conversational AI, including the ability to rewind and prune conversation hist
Renders JavaScript-heavy websites using a browser engine to fetch content for language models.
Gosub-engine is an HTML5 browser engine and web rendering pipeline that parses HTML5 and CSS3 to compute layout and render web content to pixels. It functions as a JavaScript runtime environment with a virtual machine and event loop for handling dynamic logic and asynchronous tasks. The system also includes a web storage manager for persisting cookies, local storage, and session storage. The project features a headless browser renderer capable of generating page images or extracting plain text without a visible window. It supports cross-platform graphics rendering through pluggable CPU and GP
Retrieves raw HTML and binary data from web addresses using an asynchronous networking stack with streaming support.
Cimoc is a manga reader application and cross-platform ebook viewer designed for reading digital comics and image-based documents. It functions as both an online content aggregator and an offline media library, supporting the display of media from local files and remote web sources. The application integrates various web providers through a custom parser system to fetch and display online content. It includes a synchronization system to save application settings and reading progress to a remote server, maintaining consistency across different devices. Users can customize their reading experi
Implements a pipeline that fetches remote data and processes it through parsers before passing it to the viewer.