awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
felipecsl avatar

felipecsl/wombat

0
View on GitHub↗
1,362 stars·128 forks·Ruby·MIT·4 vuesfelipecsl.github.io/wombat↗

Wombat

Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.

Features

  • Web Crawling - DSL-based web scraper for parsing structured data.
  • Ruby Crawling Frameworks - Lightweight crawler with an elegant DSL.

Historique des stars

Graphique de l'historique des stars pour felipecsl/wombatGraphique de l'historique des stars pour felipecsl/wombat

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à Wombat

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Wombat.
  • sparklemotion/mechanizeAvatar de sparklemotion

    sparklemotion/mechanize

    4,443Voir sur GitHub↗

    Mechanize is a Ruby library for web browser automation and headless browser emulation. It allows for programmatically navigating websites and simulating human behavior without a graphical user interface. The library provides an automated interface for populating and submitting web forms, including text fields, checkboxes, and file uploads. It manages stateful sessions by automatically storing and sending cookies across multiple requests to maintain user authentication and identity. Additional capabilities include web data scraping, the ability to download remote web content, and the maintena

    Ruby
    Voir sur GitHub↗4,443
  • postmodern/spidrAvatar de postmodern

    postmodern/spidr

    837Voir sur GitHub↗

    A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use.

    Ruby
    Voir sur GitHub↗837
  • propublica/uptonAvatar de propublica

    propublica/upton

    1,599Voir sur GitHub↗

    A batteries-included framework for easy web-scraping. Just add CSS! (Or do more.)

    HTML
    Voir sur GitHub↗1,599
  • lorien/web-scrapingAvatar de lorien

    lorien/web-scraping

    7,931Voir sur GitHub↗

    This project is a comprehensive resource directory for web data extraction, providing a curated collection of tools and libraries for parsing data, automating browsers, and managing network operations. It serves as a guide for extracting structured information from HTML, XML, JSON, and PDF formats. The toolkit focuses on advanced data collection strategies, including headless browser automation to interact with JavaScript and a suite of network utilities for DNS resolution and WebSocket connections. It specifically covers methods for bypassing bot protections through proxy pool management, us

    Makefile
    Voir sur GitHub↗7,931
Voir les 12 alternatives à Wombat→

Questions fréquentes

Que fait felipecsl/wombat ?

Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.

Quelles sont les fonctionnalités principales de felipecsl/wombat ?

Les fonctionnalités principales de felipecsl/wombat sont : Web Crawling, Ruby Crawling Frameworks.

Quelles sont les alternatives open-source à felipecsl/wombat ?

Les alternatives open-source à felipecsl/wombat incluent : postmodern/spidr — A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is… sparklemotion/mechanize — Mechanize is a Ruby library for web browser automation and headless browser emulation. It allows for programmatically… propublica/upton — A batteries-included framework for easy web-scraping. Just add CSS! (Or do more.). mendableai/firecrawl-mcp-server — This project is a Model Context Protocol server that connects large language models to web scraping and crawling… lorien/web-scraping — This project is a comprehensive resource directory for web data extraction, providing a curated collection of tools… adithya-s-k/omniparse — Omniparse is a multimodal content parser and generative AI ingestion engine designed to convert documents, images, and…