awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
scrapinghub avatar

scrapinghub/portia

0
View on GitHub↗
9,509 stars·1,389 forks·Python·BSD-3-Clause·15 views

Portia

Portia is a containerized scraping platform and visual web scraper that enables no-code data extraction. It serves as a Scrapy visual scraping tool and spider generator, allowing users to design and deploy web scrapers through a graphical interface instead of writing manual selector code.

The system distinguishes itself by converting visual web page annotations into executable Scrapy spider code and structured JSON specifications. This visual-to-code mapping allows users to define scraping logic and extraction rules through a point-and-click interface, which can then be exported for use in external environments.

The platform covers comprehensive web crawler management, including the ability to handle nested hierarchical data structures and monitor failed page tracking for crawl stability. It provides tools for scraping project management and extracted data retrieval, utilizing a headless browser to render pages for visual element selection.

The application is packaged for deployment via containerization to ensure consistent runtime environments.

Features

  • Visual Web Scraping Tools - Provides a graphical point-and-click interface to define scraping logic without writing code.
  • Visual Extraction Interfaces - Enables capturing structured information from web pages by annotating elements visually instead of writing scripts.
  • Hierarchical Data Structures - Organizes extracted data using parent-child relationships to capture complex page layouts like lists containing detailed items.
  • Hierarchical Extraction - Captures complex hierarchical data structures from web pages by nesting extracted items.
  • Web Data Extraction - Provides capabilities to programmatically fetch and extract structured data from specific web URLs.
  • Web Data Extraction Tools - Handles nested hierarchical structures and repetitive page patterns to extract detailed datasets from the web.
  • Spider Logic Generators - Converts visual web page annotations into executable Scrapy spider code for use in external environments.
  • Visual-to-Specification Mapping - Translates graphical UI annotations into structured JSON specifications interpreted as extraction rules by the scraping engine.
  • Headless Browsers - Integrates a headless browser to render dynamic web pages and allow visual selection of elements from the DOM.
  • Web Crawlers - Organizes multiple scraping projects and monitors page failures to ensure stable data collection workflows.
  • Web Crawling Frameworks - Leverages the Scrapy framework for high-performance asynchronous request scheduling, concurrency, and page fetching.
  • Web Scraping Management Interfaces - Ships a web-based interface for organizing, copying, and managing multiple scraping projects and spiders.
  • Project Export Environments - Converts visually defined scraping projects into executable code for deployment in external environments.
  • Container Deployment - Packages the application and its dependencies into standardized images for consistent deployment across environments.
  • Container Orchestration Deployments - Coordinates the application and its dependencies using container definitions to ensure environment consistency.
  • Containerized Scraping Platforms - Provides a web application packaged via Docker to manage and execute automated web crawls.
  • Configuration-Driven Logic - Stores scraping instructions as JSON objects instead of executable scripts to facilitate dynamic modification and export.
  • Container-Based Isolation - Utilizes container-based isolation to ensure consistent runtime environments for the scraping platform across different servers.
  • Crawl Stability Monitoring - Monitors and blocks web pages that fail to load to maintain stable and efficient crawling workflows.
  • Python Crawling Frameworks - Visual scraping tool for Scrapy.
  • Web Scraping - Visual web scraping tool.

Star history

Star history chart for scrapinghub/portiaStar history chart for scrapinghub/portia

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Portia

Similar open-source projects, ranked by how many features they share with Portia.
  • binux/pyspiderbinux avatar

    binux/pyspider

    16,809View on GitHub↗

    PySpider is a Python web crawling framework designed for automated data extraction. It provides a pipeline for periodically fetching web content, processing HTML, and persisting scraped information into database backends. The system features a web-based management interface for editing scraping scripts, monitoring task progress, and reviewing collected data. It includes a headless browser JavaScript renderer to capture rendered HTML from dynamic web pages and a distributed architecture that uses message queues to scale crawling workloads across multiple nodes. The framework also covers task

    Python
    View on GitHub↗16,809
  • ionicabizau/scrape-itIonicaBizau avatar

    IonicaBizau/scrape-it

    4,074View on GitHub↗

    scrape-it is a Node.js web scraper and HTML parser designed to extract structured data from websites and HTML files. It functions as a web data extraction tool that retrieves specific information from DOM elements and converts web content into usable data fields. The tool uses CSS selectors to target specific data points and employs schema-driven data mapping to organize unstructured web text into a consistent format. It supports custom value transformation to convert raw extracted strings into specific data formats. The system provides capabilities for web data extraction and automated cont

    JavaScripthacktoberfestnode-scraperscraper
    View on GitHub↗4,074
  • boris-code/feapderBoris-code avatar

    Boris-code/feapder

    3,709View on GitHub↗

    Feapder is a Python web crawling framework designed for building scalable data extraction systems. It features a distributed spider engine and a headless browser renderer to execute JavaScript and extract content from dynamic web pages. The system includes a scalable data deduplicator to filter duplicate URLs and records during large-scale operations. A crawler monitoring system tracks the health of active scraping jobs and triggers alerts when system anomalies occur. The framework provides capabilities for task scheduling, web data extraction, and resilient workflows that allow crawling tas

    Pythoncrawlerfeapderfeaplat
    View on GitHub↗3,709
  • ssssssss-team/spider-flowssssssss-team avatar

    ssssssss-team/spider-flow

    11,277View on GitHub↗

    Spider-flow is a Java-based web crawling and data extraction platform that provides a centralized environment for managing automated information gathering. It functions as a no-code tool, allowing users to define complex data collection pipelines through a visual, drag-and-drop interface rather than manual programming. The platform distinguishes itself through a graph-based workflow orchestration system where users link discrete nodes to define navigation and parsing logic. It supports dynamic content crawling by integrating headless browsers to execute JavaScript and render page content that

    Javacrawlerjsoupspider
    View on GitHub↗11,277
See all 30 alternatives to Portia→

Frequently asked questions

What does scrapinghub/portia do?

Portia is a containerized scraping platform and visual web scraper that enables no-code data extraction. It serves as a Scrapy visual scraping tool and spider generator, allowing users to design and deploy web scrapers through a graphical interface instead of writing manual selector code.

What are the main features of scrapinghub/portia?

The main features of scrapinghub/portia are: Visual Web Scraping Tools, Visual Extraction Interfaces, Hierarchical Data Structures, Hierarchical Extraction, Web Data Extraction, Web Data Extraction Tools, Spider Logic Generators, Visual-to-Specification Mapping.

What are some open-source alternatives to scrapinghub/portia?

Open-source alternatives to scrapinghub/portia include: binux/pyspider — PySpider is a Python web crawling framework designed for automated data extraction. It provides a pipeline for… ionicabizau/scrape-it — scrape-it is a Node.js web scraper and HTML parser designed to extract structured data from websites and HTML files.… boris-code/feapder — Feapder is a Python web crawling framework designed for building scalable data extraction systems. It features a… ssssssss-team/spider-flow — Spider-flow is a Java-based web crawling and data extraction platform that provides a centralized environment for… gsh199449/spider — Spider is a web-based platform designed for automated data extraction, providing a centralized framework to collect,… garrytan/gstack — gstack is an AI agent framework and development workflow system designed to automate the software development…