awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 مستودعات

Awesome GitHub RepositoriesHTML Document Transformation

Tools that parse and modify HTML content using selectors to dynamically update documents.

Explore 13 awesome GitHub repositories matching content management & publishing · HTML Document Transformation. Refine with filters or upvote what's useful.

Awesome HTML Document Transformation GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • oven-sh/bunالصورة الرمزية لـ oven-sh

    oven-sh/bun

    93,257عرض على GitHub↗

    Bun is a high-performance runtime environment designed to execute JavaScript and TypeScript applications with minimal latency and high throughput. Built on a native core implemented in Zig, it provides a unified execution engine that leverages JavaScriptCore for efficient memory management and low-latency startup. The project functions as an all-in-one toolchain, integrating a native bundler, transpiler, package manager, and test runner into a single command-line interface. What distinguishes Bun is its focus on native system integration and developer productivity. It features a high-performa

    Parses and modifies HTML content using CSS selectors to dynamically update documents during request or response handling.

    Rustbunbundlerjavascript
    عرض على GitHub↗93,257
  • vitejs/awesome-viteالصورة الرمزية لـ vitejs

    vitejs/awesome-vite

    16,866عرض على GitHub↗

    Awesome Vite is a curated collection of resources, plugins, and templates designed for the Vite build tool ecosystem. It serves as a central directory for developers looking to extend the capabilities of this high-performance frontend build pipeline and module bundler. The project highlights the core strengths of Vite, including its native ESM-based development server, instant hot module replacement, and pre-bundled dependency optimization. By aggregating community-maintained tools, it showcases how to leverage Vite’s plugin-based architecture to customize build pipelines, integrate popular f

    Enables programmatic modification and transformation of HTML entry files during the build process.

    JavaScriptawesomeawesome-listvite
    عرض على GitHub↗16,866
  • jhy/jsoupالصورة الرمزية لـ jhy

    jhy/jsoup

    11,340عرض على GitHub↗

    Jsoup is a Java library designed for parsing, extracting, and manipulating HTML and XML content. It provides a document object model that represents web content as a hierarchical tree, allowing for programmatic navigation and modification of elements, attributes, and text. The library functions as a toolkit for web scraping, enabling the retrieval of remote content via standard web protocols and the management of HTTP sessions for automated form interaction. The library distinguishes itself through its fault-tolerant tokenization, which reconstructs valid document structures from malformed or

    Converts raw HTML strings and streams into a structured document object model.

    Javacsscss-selectorsdom
    عرض على GitHub↗11,340
  • everywall/ladderالصورة الرمزية لـ everywall

    everywall/ladder

    8,499عرض على GitHub↗

    Ladder is a web proxy server and HTTP response modifier designed to circumvent bot protections, CORS restrictions, and paywalls. It functions by intercepting traffic to modify HTML, CSS, and JavaScript via regular expressions and altering HTTP headers to reveal restricted content. The project distinguishes itself through its ability to bypass anti-scraping mechanisms and specialized bot detection, such as Cloudflare, by integrating with external challenge-solving services. It also enables client identity emulation by spoofing user agents and network identifiers to masquerade as different brow

    Uses regular expressions to transform HTML response bodies by replacing patterns with custom content.

    Gobypasscorscors-proxy
    عرض على GitHub↗8,499
  • sparklemotion/nokogiriالصورة الرمزية لـ sparklemotion

    sparklemotion/nokogiri

    6,236عرض على GitHub↗

    Nokogiri is an XML and HTML parsing library that builds navigable document trees from strings, files, or URLs using native C parsers for speed and standards compliance. It provides a CSS selector engine that translates CSS3 selectors into XPath expressions for querying nodes, an XPath query interface with namespace support, a document manipulation toolkit for modifying parsed documents, XSD schema validation, and XSLT transformation capabilities. The library wraps libxml2 and libxslt C libraries with Ruby bindings for high-performance parsing, and integrates Google's Gumbo parser for standard

    Edit, add, or remove nodes in parsed XML and HTML documents to transform content programmatically.

    Clibxml2libxsltnokogiri
    عرض على GitHub↗6,236
  • mwilliamson/mammoth.jsالصورة الرمزية لـ mwilliamson

    mwilliamson/mammoth.js

    6,101عرض على GitHub↗

    Applies user-defined functions to modify paragraphs and runs in a .docx file's internal structure before HTML generation.

    JavaScript
    عرض على GitHub↗6,101
  • toeverything/blocksuiteالصورة الرمزية لـ toeverything

    toeverything/blocksuite

    5,544عرض على GitHub↗

    BlockSuite is a collaborative block editor framework built on a hierarchical block tree data model with CRDT-based state synchronization for real-time multi-user editing. It provides an extensible block component system, allowing developers to define custom block types through declarative schemas, services, and rendering components. The editor is packaged as cross-framework web components, making it embeddable in any JavaScript environment. The framework distinguishes itself with a command-driven editing pipeline that composes type-safe editing actions with dynamic context sharing and control

    Saves document state natively and imports/exports to standard formats via snapshot transformers.

    TypeScriptblockblock-editorcollaboration
    عرض على GitHub↗5,544
  • loro-dev/loroالصورة الرمزية لـ loro-dev

    loro-dev/loro

    5,374عرض على GitHub↗

    Loro is a conflict-free replicated data type (CRDT) framework and collaborative state engine designed for building real-time collaborative applications. It provides a distributed data synchronizer that enables multiple users to edit shared documents and complex nested structures—such as maps, lists, trees, and counters—with automatic state convergence without requiring a central server. The project distinguishes itself through a versioned document store that supports branching, forking, and merging via a directed acyclic graph of causal operation history. It enables advanced version control c

    Produces binary snapshots and incremental updates for persistent storage or synchronization across the network.

    Rustcollaborative-editingcrdtlocal-first
    عرض على GitHub↗5,374
  • apostrophecms/sanitize-htmlالصورة الرمزية لـ apostrophecms

    apostrophecms/sanitize-html

    4,129عرض على GitHub↗

    هذه مكتبة تعقيم HTML مصممة لإزالة العلامات والسمات الخطرة من HTML المقدم من المستخدم لمنع هجمات البرمجة عبر المواقع (XSS). تعمل كمرشح محتوى يقوم بإدراج عناصر وسمات محددة في القائمة البيضاء مع الهروب من الترميز غير المصرح به أو تجاهله. يتضمن المشروع محرك تحويل HTML يسمح بتعديل أو استبدال العلامات والسمات باستخدام منطق مخصص. كما يتميز بمدقق نمط CSS لتنظيف الخصائص المضمنة مقابل الأنماط المسموح بها ونظام للتحقق من صحة عنوان URL للمورد لتقييد أسماء المضيفين والمخططات. توفر المكتبة إمكانيات لتعقيم مدخلات المستخدم، والتحقق من صحة CSS المضمن، وتحويل HTML الديناميكي. يتم تنفيذ هذه الوظائف عبر محلل يدعم تصفية العلامات وتعريف العناصر المسموح بها.

    Modifies HTML content using custom logic to meet application-specific requirements during cleaning.

    JavaScript
    عرض على GitHub↗4,129
  • rstudio/bookdownالصورة الرمزية لـ rstudio

    rstudio/bookdown

    4,052عرض على GitHub↗

    Bookdown is a scientific publishing framework and multi-format document processor designed for authoring technical long-form content. It functions as an R Markdown book generator and static site generator, transforming markup files into cohesive books and reports. The system distinguishes itself through its ability to handle complex scientific document authoring, featuring integrated LaTeX typesetting, theorem environments, and automated cross-referencing for equations, figures, and theorems across multiple chapters. It enables multi-format e-book production, allowing a single project to be r

    Generates separate HTML files for each chapter while automatically updating navigation and cross-references.

    JavaScriptbookbookdownepub
    عرض على GitHub↗4,052
  • haml/hamlالصورة الرمزية لـ haml

    haml/haml

    3,834عرض على GitHub↗

    Haml is a Ruby HTML template engine and server-side rendering library. It functions as an HTML markup preprocessor that transforms a concise, indentation-based shorthand syntax into standard HTML and XHTML markup. The system uses hierarchical whitespace instead of explicit closing tags to define the structure of documents, reducing boilerplate during markup authoring. It integrates Ruby logic directly into templates to evaluate conditional statements, process dynamic data, and interpolate values. The engine provides tools for managing element attributes through hashes, controlling output whi

    Compiles a concise, indentation-based shorthand syntax into standard HTML documents.

    Ruby
    عرض على GitHub↗3,834
  • paquettg/php-html-parserالصورة الرمزية لـ paquettg

    paquettg/php-html-parser

    2,402عرض على GitHub↗

    PHP HTML Parser is a server-side programming library and DOM parser designed to ingest markup documents and structure them into a navigable parent-child object tree. The library provides a stream-based character lexer and a configurable rule engine that manages parser strictness, whitespace preservation, tag closing behavior, and character encoding detection during document ingestion. Content loading is handled through a pluggable retrieval layer that accepts local file paths, raw strings, and remote network URLs. Once loaded, documents can be queried using a chainable CSS selector engine th

    Updates node attributes, removes elements from the tree, alters text content, and manages tag properties dynamically.

    HTML
    عرض على GitHub↗2,402
  • syntax-tree/mdastالصورة الرمزية لـ syntax-tree

    syntax-tree/mdast

    1,441عرض على GitHub↗

    The project provides a standardized abstract syntax tree specification and utility library for parsing, transforming, and serializing markdown documents. It serves as a comprehensive document processing architecture that translates between raw text markup and structured node representations in both directions. The capability surface covers syntax parsing for extended markdown features, custom syntax extensions, metadata blocks, and embedded expressions, alongside document serialization that converts syntax trees back into standard and extended markup formats. It includes node modeling and val

    Translates markdown syntax trees into HTML elements or natural language structures for browser rendering and downstream processing.

    astmarkdownsyntax-tree
    عرض على GitHub↗1,441
  1. Home
  2. Content Management & Publishing
  3. Content Processing and Transformation
  4. Document Processing and Conversion
  5. Document Processing Tools
  6. Markup and Structure Parsers
  7. HTML Document Transformation

استكشف الوسوم الفرعية

  • Chapter SplittingProcesses that divide a monolithic document into separate chapter-based files while maintaining links. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: specifically focuses on splitting a single source into multiple linked HTML files.
  • Document Snapshot Persistence1 وسم فرعيSaving document state natively and importing/exporting to formats like Markdown and HTML via snapshot transformers. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: focuses on persisting and transforming the entire document state, not just modifying HTML content.
  • HTML Generation EnginesSystems that transform shorthand templates into standard HTML documents. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: focuses on the generation of HTML from templates rather than the modification of existing HTML.
  • Markdown HTML ConversionsTranslates markdown syntax tree structures into corresponding HTML markup for browser rendering. **Distinct from HTML Document Transformation:** Distinct from general HTML Document Transformation: specifically targets the translation of markdown abstract syntax trees into browser-ready HTML elements.
  • Node and Attribute EditorsProgrammatically editing, adding, or removing nodes and attributes in parsed XML and HTML documents. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: focuses on low-level node/attribute manipulation rather than selector-driven content rewriting.