awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
digininja avatar

digininja/CeWL

0
View on GitHub↗
2,575 نجوم·313 تفرعات·Ruby·5 مشاهدات

CeWL

CeWL is a custom wordlist generator and web crawling security tool designed to extract unique words and metadata from websites. It functions as an OSINT metadata extractor and security scanner, identifying potential passwords and usernames by analyzing HTML and JavaScript content.

The tool differentiates itself by combining recursive spidering with metadata extraction, allowing it to collect email addresses, author names, and creator metadata from web pages and linked files. It also captures domains, subdomains, and path components to include in generated lists.

Broad capabilities include web application spidering with depth control and regular expression filtering, as well as network request management using custom headers and proxy authentication. The system supports accessing restricted sites via Basic or Digest authentication and provides data processing utilities for word frequency analysis and list formatting.

The project is available as a containerized security scanner, packaged as a portable image to eliminate manual environment setup.

Features

  • Web Spiders - Implements a recursive web spider that traverses links to a specified depth to harvest content.
  • Custom Wordlist Generation - Spiders websites to extract unique words of a minimum length from HTML and JavaScript for use in password attacks.
  • Information Gathering - Performs reconnaissance by collecting emails and author names from website metadata.
  • Entity Extraction - Extracts email addresses and author names from mailto links and file properties for reconnaissance.
  • HTML Parsing and Extraction - Parses HTML and JavaScript to extract unique strings and words based on defined length criteria.
  • Email and Identity Extraction - Extracts email addresses and author names from web pages and linked files to build username lists.
  • Web Metadata Extractors - Extracts email addresses and author names from web responses and markup for OSINT purposes.
  • Security Crawlers - Analyzes HTML and JavaScript content to discover potential usernames and passwords.
  • Domain Structure Analyzers - Extracts domains, subdomains, and path components to analyze and catalog the hierarchical structure of a target website.
  • Crawl Boundary Controls - Uses regular expressions to define inclusion and exclusion rules to restrict automated traversal to specific domains or paths.
  • Proxy-Aware Network Clients - Supports routing network traffic through proxies with custom authentication to bypass restrictions.
  • Brute Force Attack Preparation - Prepares tailored wordlists from target domains to increase the effectiveness of dictionary attacks.
  • Containerized Scanners - Provides a portable Docker image for performing website security analysis.
  • Security Testing and Auditing - Supports security auditing by creating potential credential lists based on a company's public web presence.
  • Word Frequency Counters - Tallies word occurrences during crawling to prioritize the most frequent terms in generated lists.
  • Password and Hash Tools - Custom wordlist generator based on website content.
  • Wordlists and Enumeration - Custom wordlist generator based on website content.

سجل النجوم

مخطط تاريخ النجوم لـ digininja/cewlمخطط تاريخ النجوم لـ digininja/cewl

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ CeWL

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع CeWL.
  • andresriancho/w3afالصورة الرمزية لـ andresriancho

    andresriancho/w3af

    4,850عرض على GitHub↗

    w3af is a web penetration testing suite and security audit framework designed to identify and exploit vulnerabilities in web applications. It functions as a vulnerability scanner that crawls targets to find injection points and a fuzzer used to discover hidden endpoints and test input validation. The project distinguishes itself by providing an intercepting HTTP proxy for capturing and modifying traffic, combined with a knowledge-base driven exploitation system. It enables the execution of security exploits to gain remote shell access and supports post-exploitation activities, such as routing

    Pythonappseccross-site-scriptingscanner
    عرض على GitHub↗4,850
  • hakluke/hakrawlerالصورة الرمزية لـ hakluke

    hakluke/hakrawler

    4,993عرض على GitHub↗

    Hakrawler is a command-line web spider tool designed for security reconnaissance, built to crawl target websites and extract hyperlinks along with JavaScript file references. As a focused reconnaissance utility, it collects every discoverable URL and script source from a given domain, mapping the attack surface for penetration testing and vulnerability assessment. The tool differentiates itself through its concurrent architecture: a fixed-size goroutine pool fetches pages in parallel, while CSS selectors parse HTML to extract anchor and script references. A depth-aware recursion limiter preve

    Gobugbountycrawlinghacking
    عرض على GitHub↗4,993
  • nmap/nmapالصورة الرمزية لـ nmap

    nmap/nmap

    13,065عرض على GitHub↗

    Nmap is a command-line network security scanner and reconnaissance framework designed for infrastructure mapping and security auditing. It functions as a packet crafting utility that probes target systems to identify active hosts, detect open ports, and determine the services and operating systems running on a network. The tool distinguishes itself through its ability to perform raw socket packet injection and stateful connection tracking, allowing it to bypass standard operating system networking stacks. It utilizes an asynchronous concurrency model to manage large-scale network scans and em

    Casynchronousc-plus-pluslibpcap
    عرض على GitHub↗13,065
  • lapwinglabs/x-rayالصورة الرمزية لـ lapwinglabs

    lapwinglabs/x-ray

    5,904عرض على GitHub↗

    X-Ray is a web scraping framework and asynchronous web crawler designed to extract structured data from websites. It functions as an HTML data extractor that transforms raw page content into a defined schema using CSS-style selectors. The project implements a headless browser crawler capable of executing JavaScript to render dynamic content. It handles website content discovery through a breadth-first crawling strategy and automatic pagination discovery to traverse multi-page result sets. The framework manages web data pipelines using a concurrency-limited request queue and request rate cont

    JavaScript
    عرض على GitHub↗5,904
عرض جميع البدائل الـ 30 لـ CeWL→

الأسئلة الشائعة

ما هي وظيفة digininja/cewl؟

CeWL is a custom wordlist generator and web crawling security tool designed to extract unique words and metadata from websites. It functions as an OSINT metadata extractor and security scanner, identifying potential passwords and usernames by analyzing HTML and JavaScript content.

ما هي الميزات الرئيسية لـ digininja/cewl؟

الميزات الرئيسية لـ digininja/cewl هي: Web Spiders, Custom Wordlist Generation, Information Gathering, Entity Extraction, HTML Parsing and Extraction, Email and Identity Extraction, Web Metadata Extractors, Security Crawlers.

ما هي البدائل مفتوحة المصدر لـ digininja/cewl؟

تشمل البدائل مفتوحة المصدر لـ digininja/cewl: andresriancho/w3af — w3af is a web penetration testing suite and security audit framework designed to identify and exploit vulnerabilities… hakluke/hakrawler — Hakrawler is a command-line web spider tool designed for security reconnaissance, built to crawl target websites and… nmap/nmap — Nmap is a command-line network security scanner and reconnaissance framework designed for infrastructure mapping and… yasserg/crawler4j — Crawler4j is a multi-threaded Java web crawler and spider designed for high-volume web traversal and content… lapwinglabs/x-ray — X-Ray is a web scraping framework and asynchronous web crawler designed to extract structured data from websites. It… hect0x7/jmcomic-crawler-python — JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically…