awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
kanasimi avatar

kanasimi/work_crawler

0
View on GitHub↗
4,073 星标·352 分支·JavaScript·10 次浏览

Work Crawler

本项目是一个基于 Web 的漫画和小说下载器及多站点网络爬虫,旨在从各种媒体平台提取图像和文本。它作为一个数字媒体归档器和 EPUB 电子书生成器,使用基于插件的爬虫架构,通过特定站点的脚本定义如何从各种国际网站提取内容。

该系统通过经过身份验证的网络爬取脱颖而出,使用浏览器 Cookie 模拟来访问受限或会员专享内容。它包括用于数字漫画归档的专门功能,将图像序列组织成压缩归档文件,以及将网络小说文本和图像打包成标准化 EPUB 文件的转换流水线。

该软件提供全面的内容处理,包括确保最高质量的图像完整性验证,以及用于标准化简体和繁体中文脚本的中间件。它通过允许任务恢复的状态跟踪系统管理下载,并提供多种方式通过图形用户界面、命令行界面或 API 启动爬取。

Features

  • Digital Comic Archivers - Automates the collection of manga and webtoons from diverse platforms and organizes them into compressed image archives.
  • Plugin-Based - Uses a modular plugin-based architecture with site-specific scripts to define parsing and extraction rules for diverse platforms.
  • Automated Content Retrievers - Provides an automated system for collecting comics and novels via GUI, CLI, or API interfaces.
  • Comic and Manga Downloaders - Scrapes digital comics and novels from multiple international websites for offline reading.
  • Comic Site Extractors - Extracts images and metadata from web-based comic platforms to automate digital manga collection.
  • Incremental Novel Downloaders - Provides resumable, batch downloading of novel chapters and images for offline storage.
  • Cookie-Based Session Authentication for Downloads - Uses imported browser cookies to authenticate and download restricted or member-only content from media platforms.
  • E-book Generators - Packages extracted novel text and images into standardized electronic book files compatible with reading software.
  • Novel to EPUB Converters - Transforms long-form web novel text and images into standardized EPUB e-books.
  • Novel to EPUB Pipelines - Implements an end-to-end pipeline that scrapes web novel content and converts it directly into EPUB format.
  • Multi-Site Content Crawlers - Extracts images and text from multiple international websites using a modular system of site-specific scraping scripts.
  • Web Media Scrapers - Extracts images and text from diverse media platforms using specialized site-specific scraping logic.
  • Content Processing Pipelines - Implements a sequential workflow that fetches raw web content and processes it into structured archives and e-books.
  • Web Novel Aggregators - Implements specialized crawlers to import literary content from remote web novel catalogs.
  • Custom Scraping Logic - Implements an extensibility layer allowing users to define specific scraping rules for different websites via a crawler library.
  • Session-Cookie Persistences - Persists and reuses browser session cookies to access restricted or member-only content across multiple runs.
  • Script Conversion - Transforms text between simplified and traditional Chinese scripts to standardize the language of downloaded content.
  • Chinese Script Normalizers - Standardizes Chinese text by normalizing and converting between simplified and traditional script variants.
  • Multi-Site Title Discovery - Implements logic to search for specific titles across multiple disparate web platforms simultaneously.
  • International Content Retrievers - Retrieves digital comics and novels from international web platforms across multiple languages.
  • Sequence Archiving - Extracts image sequences from manga sites and packages them into compressed archives for local storage.
  • Multi-Platform Title Search - Locates specific titles across multiple supported websites and initiates downloads through a single action.
  • Web Content Scraping - Gathers digital books and comics from international websites by parsing HTML and network responses across different languages.
  • Manga E-book Converters - Downloads web-based literary works and converts them into e-book formats optimized for e-readers.
  • Resumable Download State Management - Tracks download checkpoints via local state files to allow resuming interrupted tasks from the last completed chapter.
  • Task Management Command-Line Interfaces - Provides a command-line interface for managing and executing download tasks with custom options and proxy settings.
  • Fetch Quality Optimizers - Ensures maximum image quality by fetching the highest available resolution and re-downloading corrupted assets.
  • Asset Quality Optimizers - Verifies the integrity of downloaded media files and automatically re-fetches corrupted assets.
  • Graphical Management Interfaces - Provides a graphical user interface for managing download configurations, themes, and multi-language settings.
  • Download Management Systems - Provides a state-tracking system to record checkpoints and resume interrupted downloads from the last completed chapter.
  • HTTP Session Simulations - Mimics authenticated browser sessions using cookies to access account-specific restricted content.
  • Multi-Interface Architectures - Provides a decoupled architecture that allows the core downloading engine to be controlled via CLI, GUI, and API.
  • Task Triggering Interfaces - Ships a unified interface to trigger and monitor the crawling and download process via graphical and command-line tools.

Star 历史

kanasimi/work_crawler 的 Star 历史图表kanasimi/work_crawler 的 Star 历史图表

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI

常见问题解答

kanasimi/work_crawler 是做什么的?

本项目是一个基于 Web 的漫画和小说下载器及多站点网络爬虫,旨在从各种媒体平台提取图像和文本。它作为一个数字媒体归档器和 EPUB 电子书生成器,使用基于插件的爬虫架构,通过特定站点的脚本定义如何从各种国际网站提取内容。

kanasimi/work_crawler 的主要功能有哪些?

kanasimi/work_crawler 的主要功能包括:Digital Comic Archivers, Plugin-Based, Automated Content Retrievers, Comic and Manga Downloaders, Comic Site Extractors, Incremental Novel Downloaders, Cookie-Based Session Authentication for Downloads, E-book Generators。

kanasimi/work_crawler 有哪些开源替代品?

kanasimi/work_crawler 的开源替代品包括: hect0x7/jmcomic-crawler-python — JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically… jack-cherish/python-spider — This is a collection of Python scripts designed for extracting data from popular Chinese websites and mobile… manga-download/hakuneko — Hakuneko is a cross-platform manga downloader and multi-platform media scraper designed to save manga and anime images… r0oth3x49/udemy-dl — udemy-dl is a Python command-line tool and web content scraper designed to download Udemy course videos, subtitles,… ripmeapp/ripme — Ripme is a batch media downloader and web media scraper designed for extracting images and videos from image-hosting… byvoid/opencc — OpenCC is a library and command-line tool for converting text between Simplified Chinese, Traditional Chinese, and…

Work Crawler 的开源替代方案

相似的开源项目,按与 Work Crawler 的功能重合度排序。
  • hect0x7/jmcomic-crawler-pythonhect0x7 的头像

    hect0x7/JMComic-Crawler-Python

    6,371在 GitHub 上查看↗

    JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically retrieve images and metadata from a comic hosting service. It functions as a media archiving tool for batch downloading albums and chapters, automating the process of saving content to a local filesystem. The project is distinguished by its ability to reverse server-side pixel obfuscation, using a decryption tool to reconstruct sliced and shuffled images. To maintain stable connectivity, it utilizes a network bypass utility featuring dynamic domain rotation and proxy routing to ci

    Python18comicasynciocrawler
    在 GitHub 上查看↗6,371
  • jack-cherish/python-spiderJack-Cherish 的头像

    Jack-Cherish/python-spider

    19,660在 GitHub 上查看↗

    This is a collection of Python scripts designed for extracting data from popular Chinese websites and mobile applications. It functions as a multi-platform data extraction toolkit, capable of automating tasks such as downloading videos from platforms like Bilibili and Douyin, scraping product reviews and images from e-commerce sites like Taobao and JD.com, and booking train tickets on the 12306 railway system. The project distinguishes itself through its focus on automating specific, high-value tasks within the Chinese internet ecosystem. It includes capabilities for solving Chinese CAPTCHA c

    Pythonpythonpython-spiderpython3
    在 GitHub 上查看↗19,660
  • manga-download/hakunekomanga-download 的头像

    manga-download/hakuneko

    6,163在 GitHub 上查看↗

    Hakuneko is a cross-platform manga downloader and multi-platform media scraper designed to save manga and anime images and videos from various websites. It functions as a tool for offline media consumption, allowing users to extract visual content from web sources and save it to local storage. The application enables cross-platform media archiving on Windows, Linux, and MacOS. It focuses on web content scraping to create local archives of images and videos, ensuring content remains accessible without an internet connection. The system manages these tasks through a connector architecture and

    JavaScriptanimeanime-downloadermanga
    在 GitHub 上查看↗6,163
  • r0oth3x49/udemy-dlr0oth3x49 的头像

    r0oth3x49/udemy-dl

    4,951在 GitHub 上查看↗

    udemy-dl is a Python command-line tool and web content scraper designed to download Udemy course videos, subtitles, and supplementary materials for offline personal use. It functions as a course media archiver that authenticates via user credentials or cookies to retrieve restricted media and metadata. The utility distinguishes itself through batch media retrieval, allowing the sequential download of multiple courses from a list of URLs. It provides granular control over the archive process, including the ability to filter specific chapters or lectures and export direct download links to a fi

    Pythoncross-platformdownload-subtitlesdownloader
    在 GitHub 上查看↗4,951
查看 Work Crawler 的所有 30 个替代方案→