awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
getmaxun avatar

getmaxun/maxun

0
View on GitHub↗
15,049 Stars·1,219 Forks·TypeScript·agpl-3.0·9 Aufrufewww.maxun.dev↗

Maxun

Maxun is an open-source web scraping and automation platform designed to transform dynamic website content into structured data. By leveraging artificial intelligence to interpret natural language prompts, the system identifies page elements and extracts information without requiring manual selector configuration. It serves as a bridge between raw web content and intelligent workflows, providing structured outputs in formats optimized for large language model ingestion and agent-based applications.

The platform distinguishes itself through its ability to handle complex, authenticated, and dynamic web environments. It synchronizes local browser sessions to access password-protected content and employs proxy rotation and browser fingerprinting to bypass anti-scraping measures. Users can orchestrate multi-step browser interactions—such as clicking buttons and filling forms—to replicate human navigation, while the self-hosted infrastructure ensures full control over data pipelines and extraction robots.

Beyond core extraction, the platform supports a broad range of automation capabilities, including recurring task scheduling, web search integration, and visual content capture. It provides programmatic access through a command-line interface and a dedicated software development kit, allowing for seamless integration with external systems via webhooks. The platform also includes monitoring tools to track website changes and distill large volumes of information into actionable insights.

Features

  • Web Scraping and Automation - Provides a platform for building and scheduling browser-based extraction workflows for AI agents.
  • Structured Data Extraction - Captures specific information from webpages and organizes it into structured formats for export or further processing.
  • Web Data Extraction - Converts raw web pages into clean, structured data formats to simplify downstream processing and automated information collection.
  • AI-Powered Web Crawlers - Uses language models to interpret web pages and transform unstructured content into structured formats.
  • Headless Browser Automation - Executes real-world browser interactions using headless engines to render dynamic content and navigate complex web interfaces.
  • Self-Hosted Infrastructure - Provides open-source infrastructure for hosting extraction robots on private servers.
  • Proxy and Fingerprint Rotation - Distributes network traffic across multiple IP addresses and fingerprint profiles to prevent detection by anti-scraping systems.
  • Session-Based Authentication Proxies - Maintains access to protected web content by synchronizing local browser session state with remote extraction workers.
  • Browser Automation Frameworks - Provides a framework for recording and replaying navigation steps to capture data from web portals.
  • Browser Automation - Records and replays browser actions like clicking buttons and filling forms to replicate manual navigation workflows.
  • AI Agent Integrations - Connects automated scraping pipelines directly to agent frameworks to supply structured data for intelligent workflows.
  • AI Agent Tool Integrations - Connects automated scraping pipelines directly to language model frameworks for intelligent processing.
  • Data Pipeline Automation - Schedules recurring data extraction tasks and integrates retrieved information into external systems via webhooks.
  • Browser Automation Orchestrators - Records and replays complex browser interactions to automate navigation across web applications.
  • Natural Language Interfaces - Uses natural language prompts to identify page elements and perform extraction tasks without manual selector configuration.
  • Browser Session Authentication - Synchronizes local browser sessions to access password-protected content without storing credentials.
  • UI Element Selectors - Uses language models to interpret natural language prompts and identify target data elements on a page.
  • Dynamic Web Scrapers - Handles complex client-side rendering and dynamic content loading to ensure consistent data extraction from modern web pages.
  • AI Automation Workflows - Integrates extraction outputs into automation platforms to enable research, data enrichment, and complex content processing.
  • AI Data Extraction - Transforms extracted web content into structured JSON, Markdown, or clean text optimized for large language model ingestion.
  • Web Search Tools - Executes search queries across the web and retrieves results as structured metadata for further processing.
  • Data Extraction Pipelines - Links extracted data to external applications through webhooks and command-line interfaces for seamless automation.
  • Web Crawlers - Navigates entire websites to extract page contents and transform them into clean markdown or HTML for downstream processing.
  • Python SDKs - Provides a dedicated software development kit to manage automated data extraction robots within custom Python codebases.
  • Webhook Integrations - Triggers external data pipelines and automated workflows by pushing extracted content to specified endpoints upon task completion.
  • Website Content Monitoring - Tracks specific website elements to detect content updates and trigger alerts.
  • Browser Task Orchestrators - Manages complex scraping workflows through a centralized control layer that schedules recurring tasks and coordinates browser interactions.
  • Browser Session Managers - Connects local browser sessions to cloud platforms to enable authenticated data extraction.
  • Data Preparation - Prepares and structures website data to facilitate the development of automated agents and data pipelines.
  • Text Summarization - Generates concise text summaries from crawled web pages to distill large volumes of information into actionable insights.
  • Automated Workflow Integration - Enables programmatic creation and management of custom scraping robots for integration into software workflows.
  • CLI Task Managers - Enables creation, execution, and monitoring of data extraction tasks directly from the terminal.
  • Screenshot Capture - Generates screenshots of web pages, including full-page captures, for visual monitoring.
  • Web APIs - Transforms dynamic website content into accessible data endpoints for consistent retrieval in automated pipelines.
  • AI Model Configurations - Allows selection between local or cloud-based language models to balance data privacy and processing performance.
  • Task Scheduling - Executes automated scraping tasks and browser workflows on a fixed timetable.

Star-Verlauf

Star-Verlauf für getmaxun/maxunStar-Verlauf für getmaxun/maxun

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Häufig gestellte Fragen

Was macht getmaxun/maxun?

Maxun is an open-source web scraping and automation platform designed to transform dynamic website content into structured data. By leveraging artificial intelligence to interpret natural language prompts, the system identifies page elements and extracts information without requiring manual selector configuration. It serves as a bridge between raw web content and intelligent workflows, providing structured outputs in formats optimized for large language model ingestion and…

Was sind die Hauptfunktionen von getmaxun/maxun?

Die Hauptfunktionen von getmaxun/maxun sind: Web Scraping and Automation, Structured Data Extraction, Web Data Extraction, AI-Powered Web Crawlers, Headless Browser Automation, Self-Hosted Infrastructure, Proxy and Fingerprint Rotation, Session-Based Authentication Proxies.

Welche Open-Source-Alternativen gibt es zu getmaxun/maxun?

Open-Source-Alternativen zu getmaxun/maxun sind unter anderem: apify/crawlee — Crawlee is a web scraping framework designed for building scalable, reliable, and distributed data extraction… apify/crawlee-python — Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive… oxylabs/ai-crawler-py — This project is an LLM-powered web crawler and data extractor that uses large language models to navigate websites and… mendableai/firecrawl — Firecrawl is a headless browser automation tool and web crawling engine designed to extract structured data from the… hangwin/mcp-chrome — This project is a Model Context Protocol tool that connects local browser instances to AI agents, enabling… automaapp/automa — Automa is a browser-based automation platform that enables users to build, schedule, and execute repetitive web tasks…

Open-Source-Alternativen zu Maxun

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Maxun.
  • apify/crawleeAvatar von apify

    apify/crawlee

    24,002Auf GitHub ansehen↗

    Crawlee is a web scraping framework designed for building scalable, reliable, and distributed data extraction pipelines. It provides a unified interface for managing headless browser automation and lightweight HTTP requests, allowing developers to handle complex web navigation, dynamic content rendering, and large-scale data collection within a single, modular architecture. The project distinguishes itself through its resource-aware concurrency controller, which dynamically scales task execution based on real-time CPU and memory usage to prevent host machine exhaustion. It also features a rob

    TypeScriptapifyautomationcrawler
    Auf GitHub ansehen↗24,002
  • apify/crawlee-pythonAvatar von apify

    apify/crawlee-python

    8,097Auf GitHub ansehen↗

    Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive tool for web scraping automation, providing a system to extract structured data from websites using both lightweight HTTP requests and headless browser automation. The framework is distinguished by its anti-bot evasion capabilities, which include browser fingerprint impersonation and tiered proxy rotation to bypass detection systems and solve challenges such as Cloudflare. It also incorporates artificial intelligence for autonomous website navigation and schema-based data extra

    Pythonapifyautomationbeautifulsoup
    Auf GitHub ansehen↗8,097
oxylabs/ai-crawler-pyAvatar von oxylabs

oxylabs/ai-crawler-py

2,683Auf GitHub ansehen↗

This project is an LLM-powered web crawler and data extractor that uses large language models to navigate websites and parse content into structured JSON or Markdown formats. It functions as an automated browser orchestrator and domain discovery engine, interpreting plain English instructions to identify relevant pages and extract specific information. The system distinguishes itself through agentic browser automation, allowing it to perform human-like interactions such as clicking buttons and scrolling based on natural language commands. It employs goal-oriented crawling to analyze website s

aiai-agentsai-crawler
Auf GitHub ansehen↗2,683
  • mendableai/firecrawlAvatar von mendableai

    mendableai/firecrawl

    139,399Auf GitHub ansehen↗

    Firecrawl is a headless browser automation tool and web crawling engine designed to extract structured data from the web. It functions as an API that transforms raw website content and documents into clean markdown and JSON formats to serve as context for large language models. The project distinguishes itself by using natural language prompts to translate human instructions into targeted data extraction tasks and browser actions. It can execute interactive page navigation, such as clicking and scrolling, and perform automated web research to retrieve structured data without manual interventi

    TypeScript
    Auf GitHub ansehen↗139,399
  • Alle 30 Alternativen zu Maxun anzeigen→