awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 रिपॉजिटरी

Awesome GitHub RepositoriesHTML Document Transformation

Tools that parse and modify HTML content using selectors to dynamically update documents.

Explore 13 awesome GitHub repositories matching content management & publishing · HTML Document Transformation. Refine with filters or upvote what's useful.

Awesome HTML Document Transformation GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • oven-sh/bunoven-sh का अवतार

    oven-sh/bun

    93,257GitHub पर देखें↗

    Bun is a high-performance runtime environment designed to execute JavaScript and TypeScript applications with minimal latency and high throughput. Built on a native core implemented in Zig, it provides a unified execution engine that leverages JavaScriptCore for efficient memory management and low-latency startup. The project functions as an all-in-one toolchain, integrating a native bundler, transpiler, package manager, and test runner into a single command-line interface. What distinguishes Bun is its focus on native system integration and developer productivity. It features a high-performa

    Parses and modifies HTML content using CSS selectors to dynamically update documents during request or response handling.

    Rustbunbundlerjavascript
    GitHub पर देखें↗93,257
  • vitejs/awesome-vitevitejs का अवतार

    vitejs/awesome-vite

    16,866GitHub पर देखें↗

    Awesome Vite is a curated collection of resources, plugins, and templates designed for the Vite build tool ecosystem. It serves as a central directory for developers looking to extend the capabilities of this high-performance frontend build pipeline and module bundler. The project highlights the core strengths of Vite, including its native ESM-based development server, instant hot module replacement, and pre-bundled dependency optimization. By aggregating community-maintained tools, it showcases how to leverage Vite’s plugin-based architecture to customize build pipelines, integrate popular f

    Enables programmatic modification and transformation of HTML entry files during the build process.

    JavaScriptawesomeawesome-listvite
    GitHub पर देखें↗16,866
  • jhy/jsoupjhy का अवतार

    jhy/jsoup

    11,340GitHub पर देखें↗

    Jsoup is a Java library designed for parsing, extracting, and manipulating HTML and XML content. It provides a document object model that represents web content as a hierarchical tree, allowing for programmatic navigation and modification of elements, attributes, and text. The library functions as a toolkit for web scraping, enabling the retrieval of remote content via standard web protocols and the management of HTTP sessions for automated form interaction. The library distinguishes itself through its fault-tolerant tokenization, which reconstructs valid document structures from malformed or

    Converts raw HTML strings and streams into a structured document object model.

    Javacsscss-selectorsdom
    GitHub पर देखें↗11,340
  • everywall/laddereverywall का अवतार

    everywall/ladder

    8,499GitHub पर देखें↗

    Ladder is a web proxy server and HTTP response modifier designed to circumvent bot protections, CORS restrictions, and paywalls. It functions by intercepting traffic to modify HTML, CSS, and JavaScript via regular expressions and altering HTTP headers to reveal restricted content. The project distinguishes itself through its ability to bypass anti-scraping mechanisms and specialized bot detection, such as Cloudflare, by integrating with external challenge-solving services. It also enables client identity emulation by spoofing user agents and network identifiers to masquerade as different brow

    Uses regular expressions to transform HTML response bodies by replacing patterns with custom content.

    Gobypasscorscors-proxy
    GitHub पर देखें↗8,499
  • sparklemotion/nokogirisparklemotion का अवतार

    sparklemotion/nokogiri

    6,236GitHub पर देखें↗

    Nokogiri is an XML and HTML parsing library that builds navigable document trees from strings, files, or URLs using native C parsers for speed and standards compliance. It provides a CSS selector engine that translates CSS3 selectors into XPath expressions for querying nodes, an XPath query interface with namespace support, a document manipulation toolkit for modifying parsed documents, XSD schema validation, and XSLT transformation capabilities. The library wraps libxml2 and libxslt C libraries with Ruby bindings for high-performance parsing, and integrates Google's Gumbo parser for standard

    Edit, add, or remove nodes in parsed XML and HTML documents to transform content programmatically.

    Clibxml2libxsltnokogiri
    GitHub पर देखें↗6,236
  • mwilliamson/mammoth.jsmwilliamson का अवतार

    mwilliamson/mammoth.js

    6,101GitHub पर देखें↗

    Applies user-defined functions to modify paragraphs and runs in a .docx file's internal structure before HTML generation.

    JavaScript
    GitHub पर देखें↗6,101
  • toeverything/blocksuitetoeverything का अवतार

    toeverything/blocksuite

    5,544GitHub पर देखें↗

    BlockSuite is a collaborative block editor framework built on a hierarchical block tree data model with CRDT-based state synchronization for real-time multi-user editing. It provides an extensible block component system, allowing developers to define custom block types through declarative schemas, services, and rendering components. The editor is packaged as cross-framework web components, making it embeddable in any JavaScript environment. The framework distinguishes itself with a command-driven editing pipeline that composes type-safe editing actions with dynamic context sharing and control

    Saves document state natively and imports/exports to standard formats via snapshot transformers.

    TypeScriptblockblock-editorcollaboration
    GitHub पर देखें↗5,544
  • loro-dev/loroloro-dev का अवतार

    loro-dev/loro

    5,374GitHub पर देखें↗

    Loro is a conflict-free replicated data type (CRDT) framework and collaborative state engine designed for building real-time collaborative applications. It provides a distributed data synchronizer that enables multiple users to edit shared documents and complex nested structures—such as maps, lists, trees, and counters—with automatic state convergence without requiring a central server. The project distinguishes itself through a versioned document store that supports branching, forking, and merging via a directed acyclic graph of causal operation history. It enables advanced version control c

    Produces binary snapshots and incremental updates for persistent storage or synchronization across the network.

    Rustcollaborative-editingcrdtlocal-first
    GitHub पर देखें↗5,374
  • apostrophecms/sanitize-htmlapostrophecms का अवतार

    apostrophecms/sanitize-html

    4,129GitHub पर देखें↗

    यह एक HTML सैनिटाइज़ेशन लाइब्रेरी है जिसे क्रॉस-साइट स्क्रिप्टिंग हमलों को रोकने के लिए उपयोगकर्ता-सबमिट किए गए HTML से खतरनाक टैग और एट्रिब्यूट को हटाने के लिए डिज़ाइन किया गया है। यह एक कंटेंट फ़िल्टर के रूप में कार्य करता है जो अनधिकृत मार्कअप को एस्केप या त्यागते समय विशिष्ट एलिमेंट और एट्रिब्यूट को व्हाइटलिस्ट करता है। प्रोजेक्ट में एक HTML परिवर्तन इंजन शामिल है जो कस्टम लॉजिक का उपयोग करके टैग और एट्रिब्यूट के संशोधन या प्रतिस्थापन की अनुमति देता है। इसमें अनुमत पैटर्न के विरुद्ध इनलाइन प्रॉपर्टी को साफ करने के लिए एक CSS स्टाइल वैलिडेटर और होस्टनाम और स्कीम को प्रतिबंधित करने के लिए संसाधन URL सत्यापन के लिए एक सिस्टम भी है। लाइब्रेरी उपयोगकर्ता इनपुट सैनिटाइज़ेशन, इनलाइन CSS सत्यापन और गतिशील HTML परिवर्तन के लिए क्षमताएं प्रदान करती है। ये फ़ंक्शन एक पार्सर के माध्यम से कार्यान्वित किए जाते हैं जो टैग फ़िल्टरिंग और अनुमत एलिमेंट की परिभाषा का समर्थन करते हैं।

    Modifies HTML content using custom logic to meet application-specific requirements during cleaning.

    JavaScript
    GitHub पर देखें↗4,129
  • rstudio/bookdownrstudio का अवतार

    rstudio/bookdown

    4,052GitHub पर देखें↗

    Bookdown एक तकनीकी प्रकाशन फ्रेमवर्क और दस्तावेज़ प्रोसेसर है जिसका उपयोग लंबी-फॉर्म प्रकाशनों, जैसे कि पुस्तकों और रिपोर्टों को लिखने के लिए किया जाता है। यह एक R Markdown पुस्तक जनरेटर और स्टेटिक साइट जनरेटर के रूप में कार्य करता है, जो उपयोगकर्ताओं को निष्पादन योग्य कोड और डेटा विज़ुअलाइज़ेशन के साथ कथा टेक्स्ट को संयोजित करने की अनुमति देता है। यह सिस्टम मल्टी-फ़ाइल असेंबली पाइपलाइन्स और कई फ़ाइलों में आंकड़ों, तालिकाओं और समीकरणों के लिए स्वचालित क्रॉस-रेफरेंस इंडेक्सिंग को प्रबंधित करने की अपनी क्षमता के माध्यम से खुद को अलग करता है। यह वैज्ञानिक सामग्री के लिए विशेष टाइपसेटिंग का समर्थन करता है, जिसमें LaTeX और HTML कंटेनरों के लिए प्रमेय और प्रमाण सिंटैक्स की मैपिंग शामिल है। यह फ्रेमवर्क PDF, EPUB, और रिस्पॉन्सिव HTML वेबसाइटों के लिए मल्टी-फॉर्मेट प्रकाशन निर्माण सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह गतिशील सामग्री एकीकरण के लिए उपकरण प्रदान करता है, जैसे कि HTML विजेट्स और इंटरैक्टिव एप्लिकेशन, साथ ही प्रोजेक्ट संरचना इनिशियलाइज़ेशन, क्लाउड होस्टिंग परिनियोजन, और सार्वजनिक कैटलॉग पंजीकरण के लिए उपयोगिताएं।

    Generates separate HTML files for each chapter while automatically updating navigation and cross-references.

    JavaScriptbookbookdownepub
    GitHub पर देखें↗4,052
  • haml/hamlhaml का अवतार

    haml/haml

    3,834GitHub पर देखें↗

    Haml is a Ruby HTML template engine and server-side rendering library. It functions as an HTML markup preprocessor that transforms a concise, indentation-based shorthand syntax into standard HTML and XHTML markup. The system uses hierarchical whitespace instead of explicit closing tags to define the structure of documents, reducing boilerplate during markup authoring. It integrates Ruby logic directly into templates to evaluate conditional statements, process dynamic data, and interpolate values. The engine provides tools for managing element attributes through hashes, controlling output whi

    Compiles a concise, indentation-based shorthand syntax into standard HTML documents.

    Ruby
    GitHub पर देखें↗3,834
  • paquettg/php-html-parserpaquettg का अवतार

    paquettg/php-html-parser

    2,402GitHub पर देखें↗

    PHP HTML Parser is a server-side programming library and DOM parser designed to ingest markup documents and structure them into a navigable parent-child object tree. The library provides a stream-based character lexer and a configurable rule engine that manages parser strictness, whitespace preservation, tag closing behavior, and character encoding detection during document ingestion. Content loading is handled through a pluggable retrieval layer that accepts local file paths, raw strings, and remote network URLs. Once loaded, documents can be queried using a chainable CSS selector engine th

    Updates node attributes, removes elements from the tree, alters text content, and manages tag properties dynamically.

    HTML
    GitHub पर देखें↗2,402
  • syntax-tree/mdastsyntax-tree का अवतार

    syntax-tree/mdast

    1,441GitHub पर देखें↗

    The project provides a standardized abstract syntax tree specification and utility library for parsing, transforming, and serializing markdown documents. It serves as a comprehensive document processing architecture that translates between raw text markup and structured node representations in both directions. The capability surface covers syntax parsing for extended markdown features, custom syntax extensions, metadata blocks, and embedded expressions, alongside document serialization that converts syntax trees back into standard and extended markup formats. It includes node modeling and val

    Translates markdown syntax trees into HTML elements or natural language structures for browser rendering and downstream processing.

    astmarkdownsyntax-tree
    GitHub पर देखें↗1,441
  1. Home
  2. Content Management & Publishing
  3. Content Processing and Transformation
  4. Document Processing and Conversion
  5. Document Processing Tools
  6. Markup and Structure Parsers
  7. HTML Document Transformation

सब-टैग एक्सप्लोर करें

  • Chapter SplittingProcesses that divide a monolithic document into separate chapter-based files while maintaining links. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: specifically focuses on splitting a single source into multiple linked HTML files.
  • Document Snapshot Persistence1 सब-टैगSaving document state natively and importing/exporting to formats like Markdown and HTML via snapshot transformers. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: focuses on persisting and transforming the entire document state, not just modifying HTML content.
  • HTML Generation EnginesSystems that transform shorthand templates into standard HTML documents. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: focuses on the generation of HTML from templates rather than the modification of existing HTML.
  • Markdown HTML ConversionsTranslates markdown syntax tree structures into corresponding HTML markup for browser rendering. **Distinct from HTML Document Transformation:** Distinct from general HTML Document Transformation: specifically targets the translation of markdown abstract syntax trees into browser-ready HTML elements.
  • Node and Attribute EditorsProgrammatically editing, adding, or removing nodes and attributes in parsed XML and HTML documents. **Distinct from HTML Document Transformation:** Distinct from HTML Document Transformation: focuses on low-level node/attribute manipulation rather than selector-driven content rewriting.