awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
lorien avatar

lorien/web-scraping

0
View on GitHub↗
7,931 stele·905 fork-uri·Makefile·8 vizualizări

Web Scraping

This project is a comprehensive resource directory for web data extraction, providing a curated collection of tools and libraries for parsing data, automating browsers, and managing network operations. It serves as a guide for extracting structured information from HTML, XML, JSON, and PDF formats.

The toolkit focuses on advanced data collection strategies, including headless browser automation to interact with JavaScript and a suite of network utilities for DNS resolution and WebSocket connections. It specifically covers methods for bypassing bot protections through proxy pool management, user agent rotation, and CAPTCHA solving integrations.

Beyond extraction, the project covers natural language processing for sentiment analysis and language detection, as well as content processing for text normalization and metadata extraction. It also includes utilities for domain analysis, geocoding services, and the management of distributed task queues for high-volume crawling.

Features

  • Web Crawling - The project crawls a domain to discover and collect information from multiple web pages.
  • Web Data Extraction - Offers a comprehensive collection of tools for programmatically scraping and structuring web content.
  • Web Crawling - Implements systems to systematically discover, navigate, and index web content at scale.
  • Natural Language Processing - Curates libraries for performing sentiment analysis, language detection, and text summarization on web content.
  • HTML and XML Parsing - Implements libraries for extracting specific data from structured HTML and XML documents.
  • Automatic Page Metadata Extraction - Provides tools for the automated retrieval of titles, meta descriptions, and social preview tags from URLs.
  • Web Page Metadata Extractors - Retrieves structured metadata and Open Graph tags from HTML markup to identify page content.
  • Headless Browser Automation - Provides tools for controlling headless browsers to interact with JavaScript and simulate users.
  • Web Scraping - Provides a curated directory of libraries, tools, and APIs for extracting and processing web data.
  • DOM Traversers - Provides algorithms for navigating and parsing the document object model to extract structured data from HTML and XML.
  • HTTP and WebSocket Clients - Provides utilities for managing network requests, DNS resolution, WebSocket connections, and proxy routing.
  • Proxy Management - Provides resources for managing proxy pools to bypass firewalls and avoid IP blocking.
  • Proxy Routing - Implements routing logic to distribute requests across rotating proxy servers to avoid IP-based restrictions.
  • JavaScript Environments - Curates runtimes and engines that execute JavaScript code outside of a standard web browser.
  • Anti-Bot Evasion - Circumvents anti-scraping protections using proxy pools, user agent rotation, and CAPTCHA solvers.
  • Automated Captcha Solvers - Integrates with external services to programmatically resolve CAPTCHA challenges during automated web navigation.
  • Protection Bypassers - Provides a directory of methods for circumventing bot protections, including IP ban avoidance and proxy management.
  • Asynchronous Task Queues - Schedules and executes high-volume scraping jobs using background queues and distributed workers.
  • Browser Automation - Provides tools for controlling headless web browsers to interact with JavaScript-rendered pages and simulate user behavior.
  • HTTP Request Clients - Offers a curated set of HTTP clients for executing programmatic requests to retrieve web pages and API data.
  • Text Similarity Scoring - Includes tools for calculating structural and token-based similarity between strings and documents using distance algorithms.
  • Data Parsing and Serialization - Curates tools for converting and extracting structured data from HTML, XML, JSON, and PDF formats.
  • Document and File Processing - Provides tools for parsing and manipulating data from CSV, JSON, and office documents.
  • Queue and Messaging - Utilizes frameworks for distributed messaging and task queues to coordinate data flow between processes.
  • URL and Network Utilities - Curates tools for parsing, cleaning, and manipulating URL components to optimize data collection.
  • Format Conversion Toolkits - Provides toolkits for programmatic transformation between diverse file formats including HTML, PDF, Markdown, and JSON.
  • Document Data Extraction - Provides utilities for parsing text, tables, and structured data from PDF and Word formats.
  • Web Article Extraction - Provides tools to isolate main body text and core content from news articles and web pages.
  • Document and Unstructured Extraction - Converts and cleans unstructured data from PDFs, CSVs, and HTML into standardized formats.
  • Text Normalization - Implements tools for normalizing Unicode and applying regular expressions to clean raw web text.
  • Text Cleaning Pipelines - Ships workflows for standardizing unstructured web text by handling contractions and formatting inconsistencies.
  • Metadata Extraction - Retrieves media, oEmbed data, and key descriptive information from web documents.
  • Background Job Processing - Provides systems for offloading and managing asynchronous tasks and child process lifecycles outside the main request flow.
  • DNS Record Resolvers - Lists utilities for performing asynchronous DNS lookups to map hostnames to IP addresses.
  • Message Brokers - Implements asynchronous communication and decoupling of background tasks using distributed message brokers.
  • Proxy Sourcing Directories - Identifies marketplaces and forums for acquiring proxy servers to maintain anonymity and avoid blocks.
  • Public Suffix Analysis - Provides resources for analyzing domain names to isolate registered domains from top-level suffixes.
  • Traffic Routing - Includes tools for implementing reverse proxies and HTTP servers to route and forward web traffic.
  • WebSocket Clients - Lists libraries for establishing bidirectional WebSocket connections for real-time data exchange.
  • Whois Lookups - Provides a directory of utilities for retrieving and parsing domain registration and ownership data via WHOIS.
  • Asynchronous I/O Libraries - Lists libraries that implement non-blocking I/O and event loops to handle high-concurrency network requests.
  • Character Encoding Utilities - Ships tools for detecting character encodings and performing internationalization and character set transformations.
  • Distributed Task Queues - Implements frameworks for distributing background work across multiple nodes using message brokers for scalable processing.
  • Robots Exclusion Compliance - Provides tools for adhering to site-specific crawling rules defined in robots.txt files.
  • User Agent Parsers - Provides a collection of tools for analyzing browser user-agent strings to identify client software metadata.
  • Web Scraping Frameworks - Collection of scraping packages across multiple programming languages.

Istoric stele

Graficul istoricului de stele pentru lorien/web-scrapingGraficul istoricului de stele pentru lorien/web-scraping

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru Web Scraping

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu Web Scraping.
  • apify/crawlee-pythonAvatar apify

    apify/crawlee-python

    8,097Vezi pe GitHub↗

    Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive tool for web scraping automation, providing a system to extract structured data from websites using both lightweight HTTP requests and headless browser automation. The framework is distinguished by its anti-bot evasion capabilities, which include browser fingerprint impersonation and tiered proxy rotation to bypass detection systems and solve challenges such as Cloudflare. It also incorporates artificial intelligence for autonomous website navigation and schema-based data extra

    Pythonapifyautomationbeautifulsoup
    Vezi pe GitHub↗8,097
  • apify/crawleeAvatar apify

    apify/crawlee

    24,002Vezi pe GitHub↗

    Crawlee is a web scraping framework designed for building scalable, reliable, and distributed data extraction pipelines. It provides a unified interface for managing headless browser automation and lightweight HTTP requests, allowing developers to handle complex web navigation, dynamic content rendering, and large-scale data collection within a single, modular architecture. The project distinguishes itself through its resource-aware concurrency controller, which dynamically scales task execution based on real-time CPU and memory usage to prevent host machine exhaustion. It also features a rob

    TypeScriptapifyautomationcrawler
    Vezi pe GitHub↗24,002
  • remitchell/python-scrapingAvatar REMitchell

    REMitchell/python-scraping

    4,714Vezi pe GitHub↗

    This project is a Python web scraping library and automated data collection suite. It provides tools for extracting structured data from websites, implementing web crawlers to navigate site links, and parsing HTML DOM structures to isolate specific elements and attributes. The toolkit includes a pipeline for processing unstructured text and cleaning raw web content to extract meaningful information. It also features capabilities for image data extraction and the integration of external APIs to retrieve structured data from remote endpoints. The system covers broad capability areas including

    Jupyter Notebook
    Vezi pe GitHub↗4,714
  • nanmicoder/crawlertutorialAvatar NanmiCoder

    NanmiCoder/CrawlerTutorial

    4,262Vezi pe GitHub↗

    CrawlerTutorial is a comprehensive Python web scraping tutorial and framework designed for extracting data from static and dynamic websites. It functions as a web data extraction pipeline and an HTTP request orchestrator, covering the full lifecycle of scraping applications from initial fetching to final data storage. The project provides specialized guidance on anti-bot bypass techniques and web API reverse engineering. It includes methods for evading browser detection through identity masking and proxy rotation, as well as techniques for identifying hidden API endpoints by analyzing network

    Python
    Vezi pe GitHub↗4,262
Vezi toate cele 30 alternative pentru Web Scraping→

Întrebări frecvente

Ce face lorien/web-scraping?

This project is a comprehensive resource directory for web data extraction, providing a curated collection of tools and libraries for parsing data, automating browsers, and managing network operations. It serves as a guide for extracting structured information from HTML, XML, JSON, and PDF formats.

Care sunt principalele funcționalități ale lorien/web-scraping?

Principalele funcționalități ale lorien/web-scraping sunt: Web Crawling, Web Data Extraction, Natural Language Processing, HTML and XML Parsing, Automatic Page Metadata Extraction, Web Page Metadata Extractors, Headless Browser Automation, Web Scraping.

Care sunt câteva alternative open-source pentru lorien/web-scraping?

Alternativele open-source pentru lorien/web-scraping includ: apify/crawlee-python — Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive… apify/crawlee — Crawlee is a web scraping framework designed for building scalable, reliable, and distributed data extraction… remitchell/python-scraping — This project is a Python web scraping library and automated data collection suite. It provides tools for extracting… nanmicoder/crawlertutorial — CrawlerTutorial is a comprehensive Python web scraping tutorial and framework designed for extracting data from static… lining0806/pythonspidernotes — PythonSpiderNotes is a comprehensive instructional resource and framework for building web crawlers and extracting… nemo2011/bilibili-api — bilibili-api is a Bilibili API wrapper and content scraper designed for programmatically accessing video metadata,…