awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
propublica avatar

propublica/upton

0
View on GitHub↗
1,599 stars·109 forks·HTML·MIT·12 views

Upton

A batteries-included framework for easy web-scraping. Just add CSS! (Or do more.)

Features

  • Web Crawling - Batteries-included framework for simplified web scraping.
  • Ruby Crawling Frameworks - Batteries-included framework for easy scraping.

Star history

Star history chart for propublica/uptonStar history chart for propublica/upton

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does propublica/upton do?

A batteries-included framework for easy web-scraping. Just add CSS! (Or do more.)

What are the main features of propublica/upton?

The main features of propublica/upton are: Web Crawling, Ruby Crawling Frameworks.

Which projects share features with propublica/upton?

Projects with overlapping indexed features include: felipecsl/wombat — Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages. sparklemotion/mechanize — Mechanize is a Ruby library for web browser automation and headless browser emulation. It allows for programmatically… postmodern/spidr — A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is… lorien/web-scraping — This project is a comprehensive resource directory for web data extraction, providing a curated collection of tools… mendableai/firecrawl-mcp-server — This project is a Model Context Protocol server that connects large language models to web scraping and crawling… adithya-s-k/omniparse — Omniparse is a multimodal content parser and generative AI ingestion engine designed to convert documents, images, and…

Projects sharing features with Upton

These projects share indexed features with Upton. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • postmodern/spidrpostmodern avatar

    postmodern/spidr

    837View on GitHub↗

    A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use.

    Ruby
    View on GitHub↗837
  • sparklemotion/mechanizesparklemotion avatar

    sparklemotion/mechanize

    4,443View on GitHub↗

    Mechanize is a Ruby library for web browser automation and headless browser emulation. It allows for programmatically navigating websites and simulating human behavior without a graphical user interface. The library provides an automated interface for populating and submitting web forms, including text fields, checkboxes, and file uploads. It manages stateful sessions by automatically storing and sending cookies across multiple requests to maintain user authentication and identity. Additional capabilities include web data scraping, the ability to download remote web content, and the maintena

    Ruby
    View on GitHub↗4,443
  • felipecsl/wombatfelipecsl avatar

    felipecsl/wombat

    1,362View on GitHub↗

    Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.

    Rubycrawlerdslruby
    View on GitHub↗1,362
  • mendableai/firecrawl-mcp-servermendableai avatar

    mendableai/firecrawl-mcp-server

    6,602View on GitHub↗

    This project is a Model Context Protocol server that connects large language models to web scraping and crawling tools. It functions as a bridge, allowing LLM clients to utilize a web crawling engine and scraping utilities to extract and process web data. The server integrates a markdown web converter that transforms dynamic web pages and PDF documents into clean markdown to optimize consumption by AI models. It also provides a browser automation interface for controlling headless sessions and bypassing access restrictions. The system covers broad capabilities including large-scale website d

    JavaScript
    View on GitHub↗6,602
Compare all 12 related projects→