For web scraper potenciado por IA para LLMs, the strongest matches are mendableai/firecrawl (Firecrawl is a headless browser automation tool designed specifically), getmaxun/maxun (Maxun is a self-hosted, open-source web scraping platform that) and firecrawl/firecrawl (Firecrawl is exactly what you need: a web data). unclecode/crawl4ai and yusufkaraaslan/skill_seekers round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Encuentra los mejores scrapers web con IA para LLMs. Compara las mejores herramientas para limpiar y estructurar datos web y encontrar la opción ideal para tu proyecto.
Firecrawl is a headless browser automation tool and web crawling engine designed to extract structured data from the web. It functions as an API that transforms raw website content and documents into clean markdown and JSON formats to serve as context for large language models. The project distinguishes itself by using natural language prompts to translate human instructions into targeted data extraction tasks and browser actions. It can execute interactive page navigation, such as clicking and scrolling, and perform automated web research to retrieve structured data without manual interventi
Firecrawl is a headless browser automation tool designed specifically to extract structured data (markdown/JSON) from web pages using natural language prompts, directly serving as context for LLMs in RAG and training workflows—hitting every key feature your search requires.
Maxun is an open-source web scraping and automation platform designed to transform dynamic website content into structured data. By leveraging artificial intelligence to interpret natural language prompts, the system identifies page elements and extracts information without requiring manual selector configuration. It serves as a bridge between raw web content and intelligent workflows, providing structured outputs in formats optimized for large language model ingestion and agent-based applications. The platform distinguishes itself through its ability to handle complex, authenticated, and dyn
Maxun is a self-hosted, open-source web scraping platform that uses AI to turn natural language prompts into structured data for LLM applications, directly matching your need for AI-driven extraction with headless browser support and API access.
Firecrawl is a web data extraction platform designed to convert unstructured web content into clean, LLM-ready formats like markdown or JSON. It functions as an autonomous web crawler and scraper, capable of mapping entire domains, performing recursive navigation, and executing complex data gathering tasks. By leveraging headless browser orchestration, the system handles dynamic, JavaScript-heavy pages to ensure comprehensive data capture. The platform distinguishes itself through its focus on agentic workflows, providing a programmatic interface that allows autonomous agents to perform live
Firecrawl is exactly what you need: a web data extraction platform that converts unstructured content into clean LLM-ready markdown or JSON, handles JavaScript-heavy pages via headless browsers, and exposes programmatic APIs for batch crawling and integration into RAG or training pipelines.
Crawl4AI is an AI-powered web crawling and data extraction engine designed to transform complex web content into structured formats. It functions as a headless browser orchestrator, enabling the navigation of dynamic websites, the execution of custom scripts, and the capture of visual assets like screenshots and PDFs. By integrating language models directly into the extraction workflow, the system converts raw HTML into clean, structured data or Markdown files optimized for downstream ingestion. The platform distinguishes itself through a distributed, self-hosted infrastructure that manages l
Crawl4AI is an AI-powered web crawling and data extraction engine that uses headless browsers and integrates language models to turn dynamic web content into clean, structured formats like JSON and Markdown, making it a perfect fit for feeding LLMs in RAG, training, or context-building pipelines.
Skill Seekers is a toolset for generating large language model knowledge bases, featuring a multi-source content scraper and a dedicated RAG data pipeline. It extracts technical data from documentation, code, and video to create structured assets and configuration files for AI-powered IDE extensions. The project distinguishes itself through the ability to transform raw data into polished tutorials and specialized skills for AI plugin marketplaces. It utilizes abstract syntax tree parsing and optical character recognition to analyze GitHub repositories, PDFs, and video frames, converting these
Skill Seekers is a multi-source content scraper and RAG data pipeline that extracts structured data from web content, documentation, code, PDFs, and video frames specifically for building LLM knowledge bases, directly matching the need for AI-driven extraction with structured output and JavaScript rendering support.
Oxylabs AI Studio Python SDK is an AI-powered web scraping toolkit that provides headless browser control, structured data extraction, and API access, directly matching the need for an LLM-ready scraping and extraction tool.
Scrapegraph-ai is a Python framework that uses large language models to automate the extraction of structured data from websites and documents. It functions as an AI-driven data extraction pipeline that converts unstructured web content into structured formats using natural language processing and graph-based logic. The project utilizes graph-based task orchestration to model scraping workflows as interconnected nodes. It features a pluggable model interface for connecting to cloud or local artificial intelligence providers and can generate executable Python code on the fly to handle site-spe
Scrapegraph-ai is a Python framework that uses LLMs to extract structured data from websites and documents, exactly matching the need for AI-powered web scraping tailored for LLM applications like RAG and context building.
Reader is an AI data ingestion pipeline and web content parser designed to convert websites and documents into clean markdown for use with large language models. It functions as a headless browser content extractor and web-to-markdown converter, transforming URLs and PDF files into structured text formats while removing irrelevant web clutter. The system optimizes retrieval augmented generation by acting as a search optimizer that retrieves web results and applies re-ranking to improve context relevance. It further enhances content accessibility by using vision models to generate descriptive
Reader is an AI-powered pipeline that converts web pages and PDFs into clean markdown using a headless browser, making it a direct fit for extracting structured web content for LLM RAG and context-building workflows, though it lacks custom extraction rules and explicit batch processing.
This project is a Model Context Protocol server that enables Large Language Models to control Playwright browsers for web automation, scraping, and end-to-end testing. It functions as a programmable interface for executing JavaScript, capturing screenshots, and interacting with web elements across multiple browser engines. The server exposes browser automation capabilities as a set of standardized tools that models can discover and invoke. It supports session-based browser isolation to ensure unique contexts for each client connection and provides a transport layer using either standard input
This MCP server lets LLMs directly control Playwright to scrape and interact with web pages, making it a fitting tool for AI-driven data extraction that can feed into RAG or training pipelines.
Browser-use is a framework for building autonomous agents that navigate, interact with, and extract data from web interfaces using natural language instructions. By acting as an orchestration layer between large language models and browser automation protocols, it enables the execution of complex, multi-step workflows without relying on brittle selectors. The system functions as a headless browser controller, providing a programmatic interface to manage browser instances and execute granular interactions. The project distinguishes itself through its ability to translate high-level intent into
Browser-use is an LLM-driven framework for browser automation and data extraction that can perform complex, multi-step scraping with structured typed output—exactly the kind of AI-powered extraction tool you need for RAG or training pipelines, though you’ll need to build the agent logic yourself.
| Repositorio | Estrellas | Lenguaje | Licencia | Último push |
|---|---|---|---|---|
| mendableai/firecrawl | 139.4K | TypeScript | AGPL-3.0 | |
| getmaxun/maxun | 15K | TypeScript | agpl-3.0 | |
| firecrawl/firecrawl | 133.5K | TypeScript | AGPL-3.0 | |
| unclecode/crawl4ai | 68.6K | Python | Apache-2.0 | |
| yusufkaraaslan/skill_seekers | 9.6K | Python | mit | |
| oxylabs/oxylabs-ai-studio-py | 2.5K | Python | mit | |
| scrapegraphai/scrapegraph-ai | 27.3K | Python | MIT | |
| jina-ai/reader | 9.8K | TypeScript | apache-2.0 | |
| executeautomation/mcp-playwright | 5.2K | TypeScript | mit | |
| browser-use/browser-use | 100.2K | Python | MIT |