awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Web scraper bazat pe LLM

Clasament actualizat la 30 iun. 2026

For web scraper bazat pe AI pentru LLM-uri, the strongest matches are mendableai/firecrawl (Firecrawl is a headless browser automation tool designed specifically), getmaxun/maxun (Maxun is a self-hosted, open-source web scraping platform that) and firecrawl/firecrawl (Firecrawl is exactly what you need: a web data). unclecode/crawl4ai and yusufkaraaslan/skill_seekers round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Găsește cele mai bune instrumente AI de web scraping pentru LLM-uri. Compară soluțiile de top pentru curățarea și structurarea datelor web pentru a găsi varianta potrivită proiectului tău.

Web scraper bazat pe LLM

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • mendableai/firecrawlAvatar mendableai

    mendableai/firecrawl

    139,399Vezi pe GitHub↗

    Firecrawl is a headless browser automation tool and web crawling engine designed to extract structured data from the web. It functions as an API that transforms raw website content and documents into clean markdown and JSON formats to serve as context for large language models. The project distinguishes itself by using natural language prompts to translate human instructions into targeted data extraction tasks and browser actions. It can execute interactive page navigation, such as clicking and scrolling, and perform automated web research to retrieve structured data without manual interventi

    Firecrawl is a headless browser automation tool designed specifically to extract structured data (markdown/JSON) from web pages using natural language prompts, directly serving as context for LLMs in RAG and training workflows—hitting every key feature your search requires.

    TypeScriptJavaScript RenderingLLM-Ready Data ExtractorsHeadless Browser Automation
    Vezi pe GitHub↗139,399
  • getmaxun/maxunAvatar getmaxun

    getmaxun/maxun

    15,049Vezi pe GitHub↗

    Maxun is an open-source web scraping and automation platform designed to transform dynamic website content into structured data. By leveraging artificial intelligence to interpret natural language prompts, the system identifies page elements and extracts information without requiring manual selector configuration. It serves as a bridge between raw web content and intelligent workflows, providing structured outputs in formats optimized for large language model ingestion and agent-based applications. The platform distinguishes itself through its ability to handle complex, authenticated, and dyn

    Maxun is a self-hosted, open-source web scraping platform that uses AI to turn natural language prompts into structured data for LLM applications, directly matching your need for AI-driven extraction with headless browser support and API access.

    TypeScriptAI Data ExtractionBrowser AutomationHeadless Browser Automation
    Vezi pe GitHub↗15,049
  • firecrawl/firecrawlAvatar firecrawl

    firecrawl/firecrawl

    133,479Vezi pe GitHub↗

    Firecrawl is a web data extraction platform designed to convert unstructured web content into clean, LLM-ready formats like markdown or JSON. It functions as an autonomous web crawler and scraper, capable of mapping entire domains, performing recursive navigation, and executing complex data gathering tasks. By leveraging headless browser orchestration, the system handles dynamic, JavaScript-heavy pages to ensure comprehensive data capture. The platform distinguishes itself through its focus on agentic workflows, providing a programmatic interface that allows autonomous agents to perform live

    Firecrawl is exactly what you need: a web data extraction platform that converts unstructured content into clean LLM-ready markdown or JSON, handles JavaScript-heavy pages via headless browsers, and exposes programmatic APIs for batch crawling and integration into RAG or training pipelines.

    TypeScriptBatch ScrapersLLM-Ready Data Extractors
    Vezi pe GitHub↗133,479
  • unclecode/crawl4aiAvatar unclecode

    unclecode/crawl4ai

    68,644Vezi pe GitHub↗

    Crawl4AI is an AI-powered web crawling and data extraction engine designed to transform complex web content into structured formats. It functions as a headless browser orchestrator, enabling the navigation of dynamic websites, the execution of custom scripts, and the capture of visual assets like screenshots and PDFs. By integrating language models directly into the extraction workflow, the system converts raw HTML into clean, structured data or Markdown files optimized for downstream ingestion. The platform distinguishes itself through a distributed, self-hosted infrastructure that manages l

    Crawl4AI is an AI-powered web crawling and data extraction engine that uses headless browsers and integrates language models to turn dynamic web content into clean, structured formats like JSON and Markdown, making it a perfect fit for feeding LLMs in RAG, training, or context-building pipelines.

    PythonHeadless
    Vezi pe GitHub↗68,644
  • yusufkaraaslan/skill_seekersAvatar yusufkaraaslan

    yusufkaraaslan/Skill_Seekers

    9,641Vezi pe GitHub↗

    Skill Seekers is a toolset for generating large language model knowledge bases, featuring a multi-source content scraper and a dedicated RAG data pipeline. It extracts technical data from documentation, code, and video to create structured assets and configuration files for AI-powered IDE extensions. The project distinguishes itself through the ability to transform raw data into polished tutorials and specialized skills for AI plugin marketplaces. It utilizes abstract syntax tree parsing and optical character recognition to analyze GitHub repositories, PDFs, and video frames, converting these

    Skill Seekers is a multi-source content scraper and RAG data pipeline that extracts structured data from web content, documentation, code, PDFs, and video frames specifically for building LLM knowledge bases, directly matching the need for AI-driven extraction with structured output and JavaScript rendering support.

    PythonJavaScript RenderingHeadless Browsers
    Vezi pe GitHub↗9,641
  • oxylabs/oxylabs-ai-studio-pyAvatar oxylabs

    oxylabs/oxylabs-ai-studio-py

    2,468Vezi pe GitHub↗

    Oxylabs AI Studio Python SDK is an AI-powered web scraping toolkit that provides headless browser control, structured data extraction, and API access, directly matching the need for an LLM-ready scraping and extraction tool.

    PythonAI-Driven Schema ExtractionsHeadless Browser Controllers
    Vezi pe GitHub↗2,468
  • scrapegraphai/scrapegraph-aiAvatar ScrapeGraphAI

    ScrapeGraphAI/Scrapegraph-ai

    27,257Vezi pe GitHub↗

    Scrapegraph-ai is a Python framework that uses large language models to automate the extraction of structured data from websites and documents. It functions as an AI-driven data extraction pipeline that converts unstructured web content into structured formats using natural language processing and graph-based logic. The project utilizes graph-based task orchestration to model scraping workflows as interconnected nodes. It features a pluggable model interface for connecting to cloud or local artificial intelligence providers and can generate executable Python code on the fly to handle site-spe

    Scrapegraph-ai is a Python framework that uses LLMs to extract structured data from websites and documents, exactly matching the need for AI-powered web scraping tailored for LLM applications like RAG and context building.

    PythonLLM-Driven Data ExtractorsAI-Powered Web CrawlersData Extraction Pipelines
    Vezi pe GitHub↗27,257
  • jina-ai/readerAvatar jina-ai

    jina-ai/reader

    9,832Vezi pe GitHub↗

    Reader is an AI data ingestion pipeline and web content parser designed to convert websites and documents into clean markdown for use with large language models. It functions as a headless browser content extractor and web-to-markdown converter, transforming URLs and PDF files into structured text formats while removing irrelevant web clutter. The system optimizes retrieval augmented generation by acting as a search optimizer that retrieves web results and applies re-ranking to improve context relevance. It further enhances content accessibility by using vision models to generate descriptive

    Reader is an AI-powered pipeline that converts web pages and PDFs into clean markdown using a headless browser, making it a direct fit for extracting structured web content for LLM RAG and context-building workflows, though it lacks custom extraction rules and explicit batch processing.

    TypeScriptHeadless Browsers
    Vezi pe GitHub↗9,832
  • executeautomation/mcp-playwrightAvatar executeautomation

    executeautomation/mcp-playwright

    5,237Vezi pe GitHub↗

    This project is a Model Context Protocol server that enables Large Language Models to control Playwright browsers for web automation, scraping, and end-to-end testing. It functions as a programmable interface for executing JavaScript, capturing screenshots, and interacting with web elements across multiple browser engines. The server exposes browser automation capabilities as a set of standardized tools that models can discover and invoke. It supports session-based browser isolation to ensure unique contexts for each client connection and provides a transport layer using either standard input

    This MCP server lets LLMs directly control Playwright to scrape and interact with web pages, making it a fitting tool for AI-driven data extraction that can feed into RAG or training pipelines.

    TypeScriptHeadless Browser Controllers
    Vezi pe GitHub↗5,237
  • browser-use/browser-useAvatar browser-use

    browser-use/browser-use

    100,229Vezi pe GitHub↗

    Browser-use is a framework for building autonomous agents that navigate, interact with, and extract data from web interfaces using natural language instructions. By acting as an orchestration layer between large language models and browser automation protocols, it enables the execution of complex, multi-step workflows without relying on brittle selectors. The system functions as a headless browser controller, providing a programmatic interface to manage browser instances and execute granular interactions. The project distinguishes itself through its ability to translate high-level intent into

    Browser-use is an LLM-driven framework for browser automation and data extraction that can perform complex, multi-step scraping with structured typed output—exactly the kind of AI-powered extraction tool you need for RAG or training pipelines, though you’ll need to build the agent logic yourself.

    PythonAutonomous Browser AgentsAutonomous Web AgentsChrome DevTools Protocols
    Vezi pe GitHub↗100,229
Compară top 10 dintr-o privire
RepositorySteleLimbajLicențăUltimul push
mendableai/firecrawl139.4KTypeScriptAGPL-3.026 iun. 2026
getmaxun/maxun15KTypeScriptagpl-3.017 feb. 2026
firecrawl/firecrawl133.5KTypeScriptAGPL-3.016 iun. 2026
unclecode/crawl4ai68.6KPythonApache-2.04 iun. 2026
yusufkaraaslan/skill_seekers9.6KPythonmit20 feb. 2026
oxylabs/oxylabs-ai-studio-py2.5KPythonmit4 dec. 2025
scrapegraphai/scrapegraph-ai27.3KPythonMIT15 iun. 2026
jina-ai/reader9.8KTypeScriptapache-2.08 mai 2025
executeautomation/mcp-playwright5.2KTypeScriptmit13 dec. 2025
browser-use/browser-use100.2KPythonMIT20 iun. 2026

Related searches

  • instrument de web scraping pentru extragerea datelor
  • a residential proxy service for web scraping
  • librărie pentru web scraping
  • un framework de web scraping pentru Python
  • un framework open source pentru web scraping
  • un crawler web de mare capacitate pentru scraping
  • framework pentru agenți AI care navighează pe web
  • AI web builder