awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
freeok avatar

freeok/so-novel

0
View on GitHub↗
7,049 星标·565 分支·Java·AGPL-3.0·8 次浏览

So Novel

so-novel is a web novel downloader and scraping engine designed to extract structured text from websites and convert it into electronic book formats. It functions as a multi-interface content extractor, providing a shared backend accessible via a web-based management dashboard, a terminal user interface, and a command line interface.

The system utilizes a rule-driven approach for data extraction, using CSS selectors and XPath rules defined in external configuration files to map web elements to specific data fields. To maintain access to content, it includes a proxy-routed request pipeline to bypass regional restrictions and anti-scraping protections.

The project converts raw extracted text into structured electronic formats such as EPUB, PDF, TXT, and HTML. It also includes utilities for translating content between Simplified and Traditional Chinese encoding standards and supports simultaneous searching across multiple aggregated web sources.

Deployment options include containerized images, shell scripts, and the ability to compile the application into a standalone native binary.

Features

  • Web Scraping - Uses customizable CSS selectors and XPath rules to extract structured content from diverse web pages.
  • Web Scraping - Extracts structured information from websites using customizable rule files based on CSS selectors, XPath, and regular expressions.
  • Artwork and Novel Downloaders - Downloads online novels and chapters from various websites for offline archival and reading.
  • Web Content Extractors - Transforms unstructured web pages into structured data formats based on predefined extraction rules.
  • Document Format Converters - Converts raw extracted text into structured electronic formats including EPUB, PDF, and TXT.
  • Document Format Conversions - Transforms extracted web text into structured electronic document formats such as EPUB, PDF, and TXT.
  • eBook Compilers - Assembles scraped web content into standardized eBook formats including EPUB, MOBI, and PDF.
  • Multi-Format Exporters - Converts extracted web content into standardized formats including EPUB, PDF, TXT, and HTML with optimized layouts.
  • Web Element Mappings - Supports mapping web elements using CSS selectors or XPaths in configuration files to identify target content.
  • Document Format Transformations - Transforms extracted raw web data into structured electronic document formats for storage and reading.
  • Pattern-Based Extraction - Defines and switches between data extraction patterns via configuration files to support diverse sources.
  • Proxy Routing - Routes outbound network traffic through configurable proxies to bypass regional restrictions and anti-scraping protections.
  • Extraction Rule Sets - Manages serialized logic used to parse and extract data from various target websites.
  • Web Element Mapping Rules - Maps website elements to specific data fields by switching between predefined configuration sets for different sources.
  • Custom Extraction Rule Definitions - Allows the creation and loading of site-specific rules that define how content is parsed from different websites.
  • CSS Selector Data Extractors - Parses unstructured web pages using configurable CSS selectors and XPath rules defined in external files.
  • Web Scraping Engines - Implements a high-performance engine for automated extraction of structured text using CSS and XPath.
  • Manga Chapter Downloaders - High-performance downloader for fetching multiple novel chapters simultaneously for offline reading.
  • System Command Dispatchers - Implements a shared backend that routes commands across web, terminal, and command line interfaces.
  • Network Proxy Configurations - Routes outbound network traffic through configurable proxy hosts to bypass regional access restrictions.
  • Multi-threaded Downloading - Fetches multiple web resources simultaneously using a configurable concurrency model and retry logic.
  • Anti-Bot Evasion - Implements techniques to bypass bot detection and security layers using proxy-routed requests.
  • Protection Bypassers - Circumvents anti-scraping security layers and bot protections by routing requests through proxy services.
  • Web Management Dashboards - Provides a browser-based dashboard for managing data extraction tasks and system configurations.
  • Multi-Interface Access - Offers a shared backend accessible via a web-based dashboard, a terminal user interface, and a command line interface.
  • Offline Web Page Archivers - Captures web-based novels and saves them as local files for offline reading on electronic devices.
  • Website Crawlers and Scrapers - Functions as a deployment-ready system for recursively retrieving and structuring web novel content.

Star 历史

freeok/so-novel 的 Star 历史图表freeok/so-novel 的 Star 历史图表

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI

So Novel 的开源替代方案

相似的开源项目,按与 So Novel 的功能重合度排序。
  • hect0x7/jmcomic-crawler-pythonhect0x7 的头像

    hect0x7/JMComic-Crawler-Python

    6,371在 GitHub 上查看↗

    JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically retrieve images and metadata from a comic hosting service. It functions as a media archiving tool for batch downloading albums and chapters, automating the process of saving content to a local filesystem. The project is distinguished by its ability to reverse server-side pixel obfuscation, using a decryption tool to reconstruct sliced and shuffled images. To maintain stable connectivity, it utilizes a network bypass utility featuring dynamic domain rotation and proxy routing to ci

    Python18comicasynciocrawler
    在 GitHub 上查看↗6,371
  • lorien/web-scrapinglorien 的头像

    lorien/web-scraping

    7,931在 GitHub 上查看↗

    This project is a comprehensive resource directory for web data extraction, providing a curated collection of tools and libraries for parsing data, automating browsers, and managing network operations. It serves as a guide for extracting structured information from HTML, XML, JSON, and PDF formats. The toolkit focuses on advanced data collection strategies, including headless browser automation to interact with JavaScript and a suite of network utilities for DNS resolution and WebSocket connections. It specifically covers methods for bypassing bot protections through proxy pool management, us

    Makefile
    在 GitHub 上查看↗7,931
  • oxylabs/how-to-scrape-amazon-product-dataoxylabs 的头像

    oxylabs/how-to-scrape-amazon-product-data

    2,511在 GitHub 上查看↗

    This project is an Amazon web scraper and e-commerce data extractor designed to retrieve product names, prices, and ratings. It functions as a headless browser crawler that converts unstructured web content from product listings into structured JSON and CSV formats. The tool incorporates anti-bot bypass capabilities to circumvent CAPTCHAs and security challenges. It achieves this through the use of residential proxy integration, automatic proxy rotation, and the modification of browser fingerprints to simulate human interaction patterns. The system provides broad web scraping capabilities, i

    amazonamazon-scraperpython
    在 GitHub 上查看↗2,511
  • gosom/google-maps-scrapergosom 的头像

    gosom/google-maps-scraper

    3,192在 GitHub 上查看↗

    This project is a distributed scraping engine designed to extract business details, customer reviews, and lead information from Google Maps. It functions as a business scraper and data extractor that can be deployed as a permanent system or as on-demand serverless functions. The system utilizes a proxy-routed web crawler to manage request origins via SOCKS5, HTTP, and HTTPS proxies. To locate contact information, it includes an email extraction tool that recursively crawls business websites linked within map listings. The software supports coordinate-based radius searches for efficient data

    Godistributed-scraperdistributed-scrapinggolang
    在 GitHub 上查看↗3,192
查看 So Novel 的所有 30 个替代方案→

常见问题解答

freeok/so-novel 是做什么的?

so-novel is a web novel downloader and scraping engine designed to extract structured text from websites and convert it into electronic book formats. It functions as a multi-interface content extractor, providing a shared backend accessible via a web-based management dashboard, a terminal user interface, and a command line interface.

freeok/so-novel 的主要功能有哪些?

freeok/so-novel 的主要功能包括:Web Scraping, Artwork and Novel Downloaders, Web Content Extractors, Document Format Converters, Document Format Conversions, eBook Compilers, Multi-Format Exporters, Web Element Mappings。

freeok/so-novel 有哪些开源替代品?

freeok/so-novel 的开源替代品包括: hect0x7/jmcomic-crawler-python — JMComic-Crawler-Python is a high-performance asynchronous web scraper and API client designed to programmatically… lorien/web-scraping — This project is a comprehensive resource directory for web data extraction, providing a curated collection of tools… oxylabs/how-to-scrape-amazon-product-data — This project is an Amazon web scraper and e-commerce data extractor designed to retrieve product names, prices, and… gosom/google-maps-scraper — This project is a distributed scraping engine designed to extract business details, customer reviews, and lead… apify/crawlee — Crawlee is a web scraping framework designed for building scalable, reliable, and distributed data extraction… asciimoo/colly — Colly is a web scraping framework and concurrent crawler written in Go. It provides a system for traversing web pages,…