awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
kanasimi avatar

kanasimi/work_crawler

0
View on GitHub↗
4,073 Stars·352 Forks·JavaScript·3 Aufrufe

Work Crawler

Dieses Projekt ist ein webbasierter Manga- und Roman-Downloader sowie ein Multi-Site-Web-Scraper, der darauf ausgelegt ist, Bilder und Texte von verschiedenen Medienplattformen zu extrahieren. Es fungiert als digitaler Medienarchivar und EPUB-Generator und nutzt eine Plugin-basierte Crawler-Architektur mit standortspezifischen Skripten, um die Extraktion von Inhalten von verschiedenen internationalen Websites zu definieren.

Das System zeichnet sich durch authentifiziertes Web-Crawling aus, wobei Browser-Cookie-Simulation verwendet wird, um auf eingeschränkte oder nur für Mitglieder zugängliche Inhalte zuzugreifen. Es enthält spezialisierte Funktionen für die Archivierung digitaler Comics, die Bildsequenzen in komprimierte Archive organisieren, sowie eine Konvertierungspipeline, die Web-Roman-Texte und Bilder in standardisierte EPUB-Dateien verpackt.

Die Software bietet eine umfassende Inhaltsverarbeitung, einschließlich der Überprüfung der Bildintegrität zur Sicherstellung maximaler Qualität sowie Middleware zur Standardisierung von vereinfachtem und traditionellem Chinesisch. Es verwaltet Downloads über ein State-Tracking-System, das die Wiederaufnahme von Aufgaben ermöglicht, und bietet mehrere Möglichkeiten, Crawls über ein grafisches Interface, ein CLI oder eine API zu starten.

Features

  • Digital Comic Archivers - Automates the collection of manga and webtoons from diverse platforms and organizes them into compressed image archives.
  • Plugin-Based - Uses a modular plugin-based architecture with site-specific scripts to define parsing and extraction rules for diverse platforms.
  • Automated Content Retrievers - Provides an automated system for collecting comics and novels via GUI, CLI, or API interfaces.
  • Comic and Manga Downloaders - Scrapes digital comics and novels from multiple international websites for offline reading.
  • Comic Site Extractors - Extracts images and metadata from web-based comic platforms to automate digital manga collection.
  • Incremental Novel Downloaders - Provides resumable, batch downloading of novel chapters and images for offline storage.
  • Cookie-Based Session Authentication for Downloads - Uses imported browser cookies to authenticate and download restricted or member-only content from media platforms.
  • E-book Generators - Packages extracted novel text and images into standardized electronic book files compatible with reading software.
  • Novel to EPUB Converters - Transforms long-form web novel text and images into standardized EPUB e-books.
  • Novel to EPUB Pipelines - Implements an end-to-end pipeline that scrapes web novel content and converts it directly into EPUB format.
  • Multi-Site Content Crawlers - Extracts images and text from multiple international websites using a modular system of site-specific scraping scripts.
  • Web Media Scrapers - Extracts images and text from diverse media platforms using specialized site-specific scraping logic.
  • Content Processing Pipelines - Implements a sequential workflow that fetches raw web content and processes it into structured archives and e-books.
  • Web Novel Aggregators - Implements specialized crawlers to import literary content from remote web novel catalogs.
  • Custom Scraping Logic - Implements an extensibility layer allowing users to define specific scraping rules for different websites via a crawler library.
  • Session-Cookie Persistences - Persists and reuses browser session cookies to access restricted or member-only content across multiple runs.
  • Script Conversion - Transforms text between simplified and traditional Chinese scripts to standardize the language of downloaded content.
  • Chinese Script Normalizers - Standardizes Chinese text by normalizing and converting between simplified and traditional script variants.
  • Multi-Site Title Discovery - Implements logic to search for specific titles across multiple disparate web platforms simultaneously.
  • International Content Retrievers - Retrieves digital comics and novels from international web platforms across multiple languages.
  • Sequence Archiving - Extracts image sequences from manga sites and packages them into compressed archives for local storage.
  • Multi-Platform Title Search - Locates specific titles across multiple supported websites and initiates downloads through a single action.
  • Web Content Scraping - Gathers digital books and comics from international websites by parsing HTML and network responses across different languages.
  • Manga E-book Converters - Downloads web-based literary works and converts them into e-book formats optimized for e-readers.
  • Resumable Download State Management - Tracks download checkpoints via local state files to allow resuming interrupted tasks from the last completed chapter.
  • Task Management Command-Line Interfaces - Provides a command-line interface for managing and executing download tasks with custom options and proxy settings.
  • Fetch Quality Optimizers - Ensures maximum image quality by fetching the highest available resolution and re-downloading corrupted assets.
  • Asset Quality Optimizers - Verifies the integrity of downloaded media files and automatically re-fetches corrupted assets.
  • Graphical Management Interfaces - Provides a graphical user interface for managing download configurations, themes, and multi-language settings.
  • Download Management Systems - Provides a state-tracking system to record checkpoints and resume interrupted downloads from the last completed chapter.
  • HTTP Session Simulations - Mimics authenticated browser sessions using cookies to access account-specific restricted content.
  • Multi-Interface Architectures - Provides a decoupled architecture that allows the core downloading engine to be controlled via CLI, GUI, and API.
  • Task Triggering Interfaces - Ships a unified interface to trigger and monitor the crawling and download process via graphical and command-line tools.

Star-Verlauf

Star-Verlauf für kanasimi/work_crawlerStar-Verlauf für kanasimi/work_crawler

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Häufig gestellte Fragen

Was macht kanasimi/work_crawler?

Dieses Projekt ist ein webbasierter Manga- und Roman-Downloader sowie ein Multi-Site-Web-Scraper, der darauf ausgelegt ist, Bilder und Texte von verschiedenen Medienplattformen zu extrahieren. Es fungiert als digitaler Medienarchivar und EPUB-Generator und nutzt eine Plugin-basierte Crawler-Architektur mit standortspezifischen Skripten, um die Extraktion von Inhalten von verschiedenen internationalen Websites zu definieren.

Was sind die Hauptfunktionen von kanasimi/work_crawler?

Die Hauptfunktionen von kanasimi/work_crawler sind: Digital Comic Archivers, Plugin-Based, Automated Content Retrievers, Comic and Manga Downloaders, Comic Site Extractors, Incremental Novel Downloaders, Cookie-Based Session Authentication for Downloads, E-book Generators.

Welche Open-Source-Alternativen gibt es zu kanasimi/work_crawler?

Open-Source-Alternativen zu kanasimi/work_crawler sind unter anderem: hect0x7/jmcomic-crawler-python — JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically… jack-cherish/python-spider — This is a collection of Python scripts designed for extracting data from popular Chinese websites and mobile… manga-download/hakuneko — Hakuneko is a cross-platform manga downloader and multi-platform media scraper designed to save manga and anime images… r0oth3x49/udemy-dl — udemy-dl is a Python command-line tool and web content scraper designed to download Udemy course videos, subtitles,… ripmeapp/ripme — Ripme is a batch media downloader and web media scraper designed for extracting images and videos from image-hosting… byvoid/opencc — OpenCC is a library and command-line tool for converting text between Simplified Chinese, Traditional Chinese, and…

Open-Source-Alternativen zu Work Crawler

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Work Crawler.
  • hect0x7/jmcomic-crawler-pythonAvatar von hect0x7

    hect0x7/JMComic-Crawler-Python

    6,371Auf GitHub ansehen↗

    JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically retrieve images and metadata from a comic hosting service. It functions as a media archiving tool for batch downloading albums and chapters, automating the process of saving content to a local filesystem. The project is distinguished by its ability to reverse server-side pixel obfuscation, using a decryption tool to reconstruct sliced and shuffled images. To maintain stable connectivity, it utilizes a network bypass utility featuring dynamic domain rotation and proxy routing to ci

    Python18comicasynciocrawler
    Auf GitHub ansehen↗6,371
  • jack-cherish/python-spiderAvatar von Jack-Cherish

    Jack-Cherish/python-spider

    19,660Auf GitHub ansehen↗

    This is a collection of Python scripts designed for extracting data from popular Chinese websites and mobile applications. It functions as a multi-platform data extraction toolkit, capable of automating tasks such as downloading videos from platforms like Bilibili and Douyin, scraping product reviews and images from e-commerce sites like Taobao and JD.com, and booking train tickets on the 12306 railway system. The project distinguishes itself through its focus on automating specific, high-value tasks within the Chinese internet ecosystem. It includes capabilities for solving Chinese CAPTCHA c

    Pythonpythonpython-spiderpython3
    Auf GitHub ansehen↗19,660
  • manga-download/hakunekoAvatar von manga-download

    manga-download/hakuneko

    6,163Auf GitHub ansehen↗

    Hakuneko is a cross-platform manga downloader and multi-platform media scraper designed to save manga and anime images and videos from various websites. It functions as a tool for offline media consumption, allowing users to extract visual content from web sources and save it to local storage. The application enables cross-platform media archiving on Windows, Linux, and MacOS. It focuses on web content scraping to create local archives of images and videos, ensuring content remains accessible without an internet connection. The system manages these tasks through a connector architecture and

    JavaScriptanimeanime-downloadermanga
    Auf GitHub ansehen↗6,163
  • r0oth3x49/udemy-dlAvatar von r0oth3x49

    r0oth3x49/udemy-dl

    4,951Auf GitHub ansehen↗

    udemy-dl is a Python command-line tool and web content scraper designed to download Udemy course videos, subtitles, and supplementary materials for offline personal use. It functions as a course media archiver that authenticates via user credentials or cookies to retrieve restricted media and metadata. The utility distinguishes itself through batch media retrieval, allowing the sequential download of multiple courses from a list of URLs. It provides granular control over the archive process, including the ability to filter specific chapters or lectures and export direct download links to a fi

    Pythoncross-platformdownload-subtitlesdownloader
    Auf GitHub ansehen↗4,951
Alle 30 Alternativen zu Work Crawler anzeigen→