awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to wspl/creeper

Open-source alternatives to Creeper

20 open-source projects similar to wspl/creeper, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Creeper alternative.

  • henrylee2cn/pholcushenrylee2cn avatar

    henrylee2cn/pholcus

    7,578View on GitHub↗

    Pholcus is a distributed web crawler framework written in Go designed for high-concurrency data extraction. It functions as a distributed crawling orchestrator and dynamic data extraction engine, utilizing a server-client architecture to coordinate tasks across multiple nodes. The system integrates a headless browser engine to render dynamic content and execute JavaScript, allowing it to extract data from single-page applications. It features a web-based management interface for configuring spider parameters and monitoring execution progress, alongside the ability to update extraction rules v

    Go
    View on GitHub↗7,578
  • montferret/ferretM

    MontFerret/ferret

    0View on GitHub↗
    View on GitHub↗0
  • puerkitobio/gocrawlPuerkitoBio avatar

    PuerkitoBio/gocrawl

    2,054View on GitHub↗

    gocrawl is a polite, slim and concurrent web crawler written in Go.

    Go
    View on GitHub↗2,054
  • crawlab-team/crawlabcrawlab-team avatar

    crawlab-team/crawlab

    12,217View on GitHub↗

    Crawlab is a distributed web scraping platform designed to centralize the management, deployment, and execution of large-scale data extraction tasks. It functions as a control plane that orchestrates scraping scripts and automated workflows across multiple nodes, providing a unified environment for managing complex data collection operations. The platform distinguishes itself through a distributed architecture that coordinates worker nodes via a central master, utilizing real-time communication to maintain oversight of all active processes. It ensures operational consistency by isolating task

    Gocrawlabcrawlercrawling-tasks
    View on GitHub↗12,217

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • geziyor/geziyorgeziyor avatar

    geziyor/geziyor

    2,773View on GitHub↗

    Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering.

    Gocrawlergoscraper
    View on GitHub↗2,773
  • gocolly/collygocolly avatar

    gocolly/colly

    25,101View on GitHub↗

    Colly is a high-performance web scraping framework designed for the automated extraction of structured data from websites. It provides a programmable toolkit that manages the complexities of large-scale data collection, including concurrent request orchestration, automatic cookie handling, and robots.txt compliance. By utilizing an asynchronous execution model, the engine maintains high throughput while preventing resource exhaustion during recursive or distributed crawling tasks. The framework is distinguished by its modular, event-driven architecture, which allows developers to hook into sp

    Gocrawlercrawlingframework
    View on GitHub↗25,101
  • hakluke/hakrawlerhakluke avatar

    hakluke/hakrawler

    4,993View on GitHub↗

    Hakrawler is a command-line web spider tool designed for security reconnaissance, built to crawl target websites and extract hyperlinks along with JavaScript file references. As a focused reconnaissance utility, it collects every discoverable URL and script source from a given domain, mapping the attack surface for penetration testing and vulnerability assessment. The tool differentiates itself through its concurrent architecture: a fixed-size goroutine pool fetches pages in parallel, while CSS selectors parse HTML to extract anchor and script references. A depth-aware recursion limiter preve

    Gobugbountycrawlinghacking
    View on GitHub↗4,993
  • jina-ai/readerjina-ai avatar

    jina-ai/reader

    9,832View on GitHub↗

    Reader is an AI data ingestion pipeline and web content parser designed to convert websites and documents into clean markdown for use with large language models. It functions as a headless browser content extractor and web-to-markdown converter, transforming URLs and PDF files into structured text formats while removing irrelevant web clutter. The system optimizes retrieval augmented generation by acting as a search optimizer that retrieves web results and applies re-ranking to improve context relevance. It further enhances content accessibility by using vision models to generate descriptive

    TypeScriptllmproxy
    View on GitHub↗9,832
  • mendableai/firecrawlmendableai avatar

    mendableai/firecrawl

    139,399View on GitHub↗

    Firecrawl is a headless browser automation tool and web crawling engine designed to extract structured data from the web. It functions as an API that transforms raw website content and documents into clean markdown and JSON formats to serve as context for large language models. The project distinguishes itself by using natural language prompts to translate human instructions into targeted data extraction tasks and browser actions. It can execute interactive page navigation, such as clicking and scrolling, and perform automated web research to retrieve structured data without manual interventi

    TypeScript
    View on GitHub↗139,399
  • projectdiscovery/katanaprojectdiscovery avatar

    projectdiscovery/katana

    15,584View on GitHub↗

    Katana is a web crawler and spider designed for security reconnaissance and web application mapping. It functions as a utility for identifying endpoints, forms, and API structures across web targets by combining standard HTTP request traversal with headless browser automation to render dynamic, JavaScript-heavy content. The tool distinguishes itself through its ability to maintain authenticated sessions and handle complex web interactions, such as automated form submission and captcha resolution. It provides granular control over the discovery process, allowing users to define specific crawl

    Goclicrawlergocrawler
    View on GitHub↗15,584
  • puerkitobio/fetchbotPuerkitoBio avatar

    PuerkitoBio/fetchbot

    791View on GitHub↗

    Package fetchbot provides a simple and flexible web crawler that follows the robots.txt policies and crawl delays.

    Go
    View on GitHub↗791
  • raviqqe/muffetraviqqe avatar

    raviqqe/muffet

    2,612View on GitHub↗

    Fast website link checker in Go

    Go
    View on GitHub↗2,612
  • shiyanhui/dhtshiyanhui avatar

    shiyanhui/dht

    2,774View on GitHub↗

    See the video on the Youtube.

    Go
    View on GitHub↗2,774
  • slotix/dataflowkitslotix avatar

    slotix/dataflowkit

    714View on GitHub↗

    Dataflow kit ("DFK") is a Web Scraping framework for Gophers. It extracts data from web pages, following the specified CSS Selectors.

    Go
    View on GitHub↗714
  • unclecode/crawl4aiunclecode avatar

    unclecode/crawl4ai

    68,644View on GitHub↗

    Crawl4AI is an AI-powered web crawling and data extraction engine designed to transform complex web content into structured formats. It functions as a headless browser orchestrator, enabling the navigation of dynamic websites, the execution of custom scripts, and the capture of visual assets like screenshots and PDFs. By integrating language models directly into the extraction workflow, the system converts raw HTML into clean, structured data or Markdown files optimized for downstream ingestion. The platform distinguishes itself through a distributed, self-hosted infrastructure that manages l

    Python
    View on GitHub↗68,644
  • wcong/ants-gowcong avatar

    wcong/ants-go

    361View on GitHub↗

    open source, restful, distributed crawler engine

    Go
    View on GitHub↗361
  • amirgamil/apolloamirgamil avatar

    amirgamil/apollo

    1,376View on GitHub↗

    A Unix-style personal search engine and web crawler for your digital footprint.

    Gopersonal-searchposeidonsearch
    View on GitHub↗1,376
  • yhat/scrapeyhat avatar

    yhat/scrape

    1,515View on GitHub↗

    A simple, higher level interface for Go web scraping.

    Go
    View on GitHub↗1,515
  • antchfx/antchantchfx avatar

    antchfx/antch

    267View on GitHub↗

    Antch, a fast, powerful and extensible web crawling & scraping framework for Go

    Go
    View on GitHub↗267
  • asciimoo/collyasciimoo avatar

    asciimoo/colly

    25,348View on GitHub↗

    Colly is a web scraping framework and concurrent crawler written in Go. It provides a system for traversing web pages, following links, and extracting structured data from HTML and XML documents. The framework includes a distributed scraping engine designed to spread data collection tasks across multiple instances to increase throughput. It ensures compliance with website owner policies by automatically reading and respecting robots.txt files. The system manages request lifecycles through domain-based rate limiting, concurrency controls, and session management via a stateful cookie jar. It s

    Go
    View on GitHub↗25,348