awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
mishushakov avatar

mishushakov/llm-scraper

0
View on GitHub↗
6,190 Stars·369 Forks·TypeScript·mit·6 Aufrufe

Llm Scraper

Features

  • LLM-Powered Scrapers - A tool that uses a language model to extract structured data from web pages based on a user-defined schema.
  • Schema-Driven Extraction - Defines the shape of data to extract from webpages with Zod or JSON schemas.
  • Webpage Parsers - Uses a language model to parse raw webpage content and return structured data matching a predefined schema.
  • Structured Data Extraction - Parse webpage content with a language model to return fields defined by a Zod or JSON schema.
  • LLM-Powered Webpage Parsers - Uses an LLM to parse a page's content and return typed fields defined by a Zod or JSON Schema.
  • Headless Browser Automation - Leverages Playwright to programmatically control a browser for automated web content extraction.
  • LLM-Powered Scrapers - Extracts structured data from web pages using a language model and a user-defined schema.
  • Incremental Result Streaming - Yields extracted objects incrementally as the language model generates them for early consumption.
  • Extraction Streams - Receives partial structured data incrementally as the language model processes a webpage.
  • Provider-Agnostic Interfaces - Designed to work with various language models, allowing users to choose the provider or model.
  • Zod Schemas - Accepts Zod schemas directly to define the structure and validation rules for extracted data fields.
  • Schema Definitions - Supports JSON Schema as an alternative to Zod for defining extraction schemas.
  • Incremental Result Streams - Yield partial structured results incrementally as the language model produces them for early consumption.
  • Incremental Streams - Yields partial structured results incrementally as the language model produces them during extraction.
  • Playwright Scripts - Produces a standalone Playwright script that replicates the extraction logic without requiring an LLM at runtime.
  • Playwright Script Generators - Produces a Playwright script that extracts the same schema-defined data without an LLM call.
  • Playwright Scripts - Creates reusable Playwright scripts that extract data without further language model calls.

Star-Verlauf

Star-Verlauf für mishushakov/llm-scraperStar-Verlauf für mishushakov/llm-scraper

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Llm Scraper

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Llm Scraper.
  • itsowen/cyberscraper-2077Avatar von itsOwen

    itsOwen/CyberScraper-2077

    2,887Auf GitHub ansehen↗

    CyberScraper-2077 is an AI-powered web scraping tool that uses large language models to extract and structure data from websites into organized formats. It functions as an LLM web scraper and AI content parser, transforming unstructured raw web text into specific data schemas. The project distinguishes itself through a suite of anonymity and evasion tools, including proxy rotation, SOCKS-based identity masking, and the ability to route traffic through the Tor network to access hidden onion services. It further includes a bot detection bypass system that employs stealth parameters and custom n

    Pythonai-scrapinggemini-apillm
    Auf GitHub ansehen↗2,887
  • apify/crawlee-pythonAvatar von apify

    apify/crawlee-python

    8,097Auf GitHub ansehen↗

    Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive tool for web scraping automation, providing a system to extract structured data from websites using both lightweight HTTP requests and headless browser automation. The framework is distinguished by its anti-bot evasion capabilities, which include browser fingerprint impersonation and tiered proxy rotation to bypass detection systems and solve challenges such as Cloudflare. It also incorporates artificial intelligence for autonomous website navigation and schema-based data extra

    Pythonapifyautomationbeautifulsoup
    Auf GitHub ansehen↗8,097
  • getmaxun/maxunAvatar von getmaxun

    getmaxun/maxun

    15,049Auf GitHub ansehen↗

    Maxun is an open-source web scraping and automation platform designed to transform dynamic website content into structured data. By leveraging artificial intelligence to interpret natural language prompts, the system identifies page elements and extracts information without requiring manual selector configuration. It serves as a bridge between raw web content and intelligent workflows, providing structured outputs in formats optimized for large language model ingestion and agent-based applications. The platform distinguishes itself through its ability to handle complex, authenticated, and dyn

    TypeScriptagentsapiautomation
    Auf GitHub ansehen↗15,049
  • any4ai/anycrawlAvatar von any4ai

    any4ai/AnyCrawl

    2,742Auf GitHub ansehen↗

    AnyCrawl is an AI-powered data extractor, automated web crawler, and headless browser orchestrator. It serves as a web content extraction API and a gateway that connects crawling and scraping tools to language models using a standardized API protocol. The project specializes in converting unstructured website content into structured JSON or markdown optimized for AI assistants. It utilizes language models and JSON schemas to pull specific information into validated formats and provides capabilities for AI page summarization and LLM-optimized content extraction. The system manages comprehensi

    TypeScriptai-scrapingaitoolscrawl
    Auf GitHub ansehen↗2,742
Alle 30 Alternativen zu Llm Scraper anzeigen→

Häufig gestellte Fragen

Was sind die Hauptfunktionen von mishushakov/llm-scraper?

Die Hauptfunktionen von mishushakov/llm-scraper sind: LLM-Powered Scrapers, Schema-Driven Extraction, Webpage Parsers, Structured Data Extraction, LLM-Powered Webpage Parsers, Headless Browser Automation, Incremental Result Streaming, Extraction Streams.

Welche Open-Source-Alternativen gibt es zu mishushakov/llm-scraper?

Open-Source-Alternativen zu mishushakov/llm-scraper sind unter anderem: itsowen/cyberscraper-2077 — CyberScraper-2077 is an AI-powered web scraping tool that uses large language models to extract and structure data… apify/crawlee-python — Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive… getmaxun/maxun — Maxun is an open-source web scraping and automation platform designed to transform dynamic website content into… browserbase/mcp-server-browserbase — This project is an MCP browser automation server that connects large language models to headless cloud browsers. It… axa-group/parsr — Parsr is an unstructured data extractor and document parsing pipeline that converts raw files and images into cleaned,… any4ai/anycrawl — AnyCrawl is an AI-powered data extractor, automated web crawler, and headless browser orchestrator. It serves as a web…