awesome-repositories.com
Blog
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectAboutHow we rankPressMCP server
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
kanasimi avatar

kanasimi/work_crawler

0
View on GitHub↗
4,073 stars·352 forks·JavaScript·3 views

Work Crawler

This project is a web-based manga and novel downloader and multi-site web scraper designed to extract images and text from diverse media platforms. It functions as a digital media archiver and EPUB e-book generator, using a plugin-based crawler architecture with site-specific scripts to define how content is extracted from various international websites.

The system distinguishes itself through authenticated web crawling, using browser cookie simulation to access restricted or member-only content. It includes specialized capabilities for digital comic archiving, which organizes image sequences into compressed archives, and a conversion pipeline that packages web novel text and images into standardized EPUB files.

The software provides comprehensive content processing, including image integrity verification to ensure maximum quality and middleware for standardizing simplified and traditional Chinese scripts. It manages downloads via a state-tracking system that allows for task resumption, and offers multiple ways to initiate crawls through a graphical user interface, command-line interface, or API.

Features

  • Digital Comic Archivers - Automates the collection of manga and webtoons from diverse platforms and organizes them into compressed image archives.
  • Plugin-Based - Uses a modular plugin-based architecture with site-specific scripts to define parsing and extraction rules for diverse platforms.
  • Automated Content Retrievers - Provides an automated system for collecting comics and novels via GUI, CLI, or API interfaces.
  • Comic and Manga Downloaders - Scrapes digital comics and novels from multiple international websites for offline reading.
  • Comic Site Extractors - Extracts images and metadata from web-based comic platforms to automate digital manga collection.
  • Incremental Novel Downloaders - Provides resumable, batch downloading of novel chapters and images for offline storage.
  • Cookie-Based Session Authentication for Downloads - Uses imported browser cookies to authenticate and download restricted or member-only content from media platforms.
  • E-book Generators - Packages extracted novel text and images into standardized electronic book files compatible with reading software.
  • Novel to EPUB Converters - Transforms long-form web novel text and images into standardized EPUB e-books.
  • Novel to EPUB Pipelines - Implements an end-to-end pipeline that scrapes web novel content and converts it directly into EPUB format.
  • Multi-Site Content Crawlers - Extracts images and text from multiple international websites using a modular system of site-specific scraping scripts.
  • Web Media Scrapers - Extracts images and text from diverse media platforms using specialized site-specific scraping logic.
  • Content Processing Pipelines - Implements a sequential workflow that fetches raw web content and processes it into structured archives and e-books.
  • Web Novel Aggregators - Implements specialized crawlers to import literary content from remote web novel catalogs.
  • Custom Scraping Logic - Implements an extensibility layer allowing users to define specific scraping rules for different websites via a crawler library.
  • Session-Cookie Persistences - Persists and reuses browser session cookies to access restricted or member-only content across multiple runs.
  • Script Conversion - Transforms text between simplified and traditional Chinese scripts to standardize the language of downloaded content.
  • Chinese Script Normalizers - Standardizes Chinese text by normalizing and converting between simplified and traditional script variants.
  • Multi-Site Title Discovery - Implements logic to search for specific titles across multiple disparate web platforms simultaneously.
  • International Content Retrievers - Retrieves digital comics and novels from international web platforms across multiple languages.
  • Sequence Archiving - Extracts image sequences from manga sites and packages them into compressed archives for local storage.
  • Multi-Platform Title Search - Locates specific titles across multiple supported websites and initiates downloads through a single action.
  • Web Content Scraping - Gathers digital books and comics from international websites by parsing HTML and network responses across different languages.
  • Manga E-book Converters - Downloads web-based literary works and converts them into e-book formats optimized for e-readers.
  • Resumable Download State Management - Tracks download checkpoints via local state files to allow resuming interrupted tasks from the last completed chapter.
  • Task Management Command-Line Interfaces - Provides a command-line interface for managing and executing download tasks with custom options and proxy settings.
  • Fetch Quality Optimizers - Ensures maximum image quality by fetching the highest available resolution and re-downloading corrupted assets.
  • Asset Quality Optimizers - Verifies the integrity of downloaded media files and automatically re-fetches corrupted assets.
  • Graphical Management Interfaces - Provides a graphical user interface for managing download configurations, themes, and multi-language settings.
  • Download Management Systems - Provides a state-tracking system to record checkpoints and resume interrupted downloads from the last completed chapter.
  • HTTP Session Simulations - Mimics authenticated browser sessions using cookies to access account-specific restricted content.
  • Multi-Interface Architectures - Provides a decoupled architecture that allows the core downloading engine to be controlled via CLI, GUI, and API.
  • Task Triggering Interfaces - Ships a unified interface to trigger and monitor the crawling and download process via graphical and command-line tools.

Star history

Star history chart for kanasimi/work_crawlerStar history chart for kanasimi/work_crawler

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Work Crawler

Similar open-source projects, ranked by how many features they share with Work Crawler.
  • hect0x7/jmcomic-crawler-pythonhect0x7 avatar

    hect0x7/JMComic-Crawler-Python

    6,371View on GitHub↗

    JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically retrieve images and metadata from a comic hosting service. It functions as a media archiving tool for batch downloading albums and chapters, automating the process of saving content to a local filesystem. The project is distinguished by its ability to reverse server-side pixel obfuscation, using a decryption tool to reconstruct sliced and shuffled images. To maintain stable connectivity, it utilizes a network bypass utility featuring dynamic domain rotation and proxy routing to ci

    Python18comicasynciocrawler
    View on GitHub↗6,371
  • jack-cherish/python-spiderJack-Cherish avatar

    Jack-Cherish/python-spider

    19,660View on GitHub↗

    This is a collection of Python scripts designed for extracting data from popular Chinese websites and mobile applications. It functions as a multi-platform data extraction toolkit, capable of automating tasks such as downloading videos from platforms like Bilibili and Douyin, scraping product reviews and images from e-commerce sites like Taobao and JD.com, and booking train tickets on the 12306 railway system. The project distinguishes itself through its focus on automating specific, high-value tasks within the Chinese internet ecosystem. It includes capabilities for solving Chinese CAPTCHA c

    Pythonpythonpython-spiderpython3
    View on GitHub↗19,660
  • manga-download/hakunekomanga-download avatar

    manga-download/hakuneko

    6,163View on GitHub↗

    Hakuneko is a cross-platform manga downloader and multi-platform media scraper designed to save manga and anime images and videos from various websites. It functions as a tool for offline media consumption, allowing users to extract visual content from web sources and save it to local storage. The application enables cross-platform media archiving on Windows, Linux, and MacOS. It focuses on web content scraping to create local archives of images and videos, ensuring content remains accessible without an internet connection. The system manages these tasks through a connector architecture and

    JavaScriptanimeanime-downloadermanga
    View on GitHub↗6,163
  • r0oth3x49/udemy-dlr0oth3x49 avatar

    r0oth3x49/udemy-dl

    4,951View on GitHub↗

    udemy-dl is a Python command-line tool and web content scraper designed to download Udemy course videos, subtitles, and supplementary materials for offline personal use. It functions as a course media archiver that authenticates via user credentials or cookies to retrieve restricted media and metadata. The utility distinguishes itself through batch media retrieval, allowing the sequential download of multiple courses from a list of URLs. It provides granular control over the archive process, including the ability to filter specific chapters or lectures and export direct download links to a fi

    Pythoncross-platformdownload-subtitlesdownloader
    View on GitHub↗4,951
See all 30 alternatives to Work Crawler→

Frequently asked questions

What does kanasimi/work_crawler do?

This project is a web-based manga and novel downloader and multi-site web scraper designed to extract images and text from diverse media platforms. It functions as a digital media archiver and EPUB e-book generator, using a plugin-based crawler architecture with site-specific scripts to define how content is extracted from various international websites.

What are the main features of kanasimi/work_crawler?

The main features of kanasimi/work_crawler are: Digital Comic Archivers, Plugin-Based, Automated Content Retrievers, Comic and Manga Downloaders, Comic Site Extractors, Incremental Novel Downloaders, Cookie-Based Session Authentication for Downloads, E-book Generators.

What are some open-source alternatives to kanasimi/work_crawler?

Open-source alternatives to kanasimi/work_crawler include: hect0x7/jmcomic-crawler-python — JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically… jack-cherish/python-spider — This is a collection of Python scripts designed for extracting data from popular Chinese websites and mobile… manga-download/hakuneko — Hakuneko is a cross-platform manga downloader and multi-platform media scraper designed to save manga and anime images… r0oth3x49/udemy-dl — udemy-dl is a Python command-line tool and web content scraper designed to download Udemy course videos, subtitles,… ripmeapp/ripme — Ripme is a batch media downloader and web media scraper designed for extracting images and videos from image-hosting… byvoid/opencc — OpenCC is a library and command-line tool for converting text between Simplified Chinese, Traditional Chinese, and…