awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
coolwanglu avatar

coolwanglu/pdf2htmlEXArchived

0
View on GitHub↗
10,603 نجوم·1,844 تفرعات·HTML·6 مشاهداتcoolwanglu.github.com/pdf2htmlEX↗

Pdf2htmlEX

pdf2htmlEX is a tool that converts PDF documents into HTML while preserving the original text, fonts, and layout. It uses CSS positioning and font embedding to replicate the PDF's appearance in a browser, producing output that works without JavaScript. The tool can generate a single self-contained HTML file with all resources embedded, or split the document into separate HTML files per page for individual loading and navigation.

The converter offers extensive control over the output, including the ability to embed fonts directly into the HTML using base64-encoded Data URIs, or keep them as separate files for caching. It supports page range selection, output location configuration, and image fallback rendering when vector conversion fails. The tool also provides options for custom CSS overrides, template customization, and resource embedding control to balance file size against HTTP requests.

Additional capabilities include font metadata inspection, duplicate font optimization, and font size precision maintenance. The output can preserve hyperlinks, bookmarks, and print functionality from the original PDF, and supports vertical writing mode for certain text layouts. For deployment, the tool can be run in a Docker container and supports HTTP compression and mobile optimization.

Features

  • PDF to HTML Converters - Renders PDF files as single HTML documents, preserving text, fonts, layout, and embedded elements.
  • Self-Contained HTML Reports - Generates single-file HTML outputs with embedded fonts and images for portable distribution.
  • PDF Layout Preservers - Maintains exact font sizes, styles, and text positioning from PDF in the HTML output.
  • Page-to-File Splitting - Splits PDFs into separate HTML files per page for individual loading and navigation.
  • Semantic HTML Structures - Converts PDF documents into structured HTML that preserves text, fonts, and layout for web display.
  • PDF Interactive Elements - Preserves hyperlinks, bookmarks, and print functionality from PDFs in HTML output.
  • Data URI Embeddings - Embeds fonts and images as data URIs in HTML for self-contained documents.
  • Font Data URI Embedders - Embeds fonts as base64 Data URIs to preserve original typography without external dependencies.
  • PDF Layout Replicators - Provides CSS-based layout replication to preserve exact PDF page geometry in HTML output.
  • PDF Font Embedding - Embeds fonts from PDF files directly into HTML output to maintain original typography.
  • PDF Layout Preservation Patterns - Uses CSS positioning and font embedding to replicate PDF layout and text flow in the browser.
  • JavaScript-Free Web Interfaces - Generates fully functional HTML output that displays correctly without JavaScript, relying solely on CSS.
  • HTML Page Decomposition - Splits PDF into separate HTML files per page for lazy loading and dynamic navigation via AJAX.
  • Document Output Appearance Configurations - Adjusts the visual style of the generated HTML through configuration options for fonts, colors, and layout.
  • PDF Navigational Bookmarks - Retains hyperlinks and table-of-contents outlines from PDFs for navigation in HTML.
  • Self-Contained Document Exports - Embeds all resources like fonts and images into a single HTML file for portable distribution.
  • Dimension Adjustments - Sets zoom factor or maximum page dimensions to scale rendered HTML output.
  • Page Layout Adjustments - Adjusts page scaling to accommodate differences between PDF and HTML formats.
  • Document Page Loaders - Loads individual pages on scroll to reduce initial download size and wait time.
  • PDF Interactive Feature Preservers - Preserves hyperlinks, bookmarks, and page structure from the original PDF for interactive browsing.
  • PDF-to-HTML Performance Optimizations - Optimizes PDF-to-HTML output for performance, caching, and mobile rendering.
  • Batch Document Processing - Converts multiple PDF pages or documents with configurable output settings and resource management.
  • Resource Separation - Stores fonts, images, CSS, and JavaScript in external files for browser caching.
  • Browser Cache Optimizers - Outputs CSS, fonts, and images as separate files for independent browser caching and reduced HTTP requests.
  • CSS Styling - Allows overriding default styles with custom CSS to control the visual appearance of converted output.
  • Conversion Output Template Customizers - Modifies the HTML, CSS, and JavaScript templates that control how the PDF content is rendered and styled in the browser.
  • Conversion Output CSS Overrides - Includes custom CSS after the default styles to override the visual appearance of elements extracted from the PDF.
  • Output Zoom Setters - Sets a zoom factor to scale rendered page dimensions in the output HTML.
  • PDF Rendering Configurators - Configures zoom, page range, and output width to control PDF rendering in HTML.
  • Conversion Resource Embedding Controls - Chooses which elements like CSS, fonts, images, JavaScript, and outlines to embed in the HTML or keep as separate files.

سجل النجوم

مخطط تاريخ النجوم لـ coolwanglu/pdf2htmlexمخطط تاريخ النجوم لـ coolwanglu/pdf2htmlex

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة coolwanglu/pdf2htmlex؟

pdf2htmlEX is a tool that converts PDF documents into HTML while preserving the original text, fonts, and layout. It uses CSS positioning and font embedding to replicate the PDF's appearance in a browser, producing output that works without JavaScript. The tool can generate a single self-contained HTML file with all resources embedded, or split the document into separate HTML files per page for individual loading and navigation.

ما هي الميزات الرئيسية لـ coolwanglu/pdf2htmlex؟

الميزات الرئيسية لـ coolwanglu/pdf2htmlex هي: PDF to HTML Converters, Self-Contained HTML Reports, PDF Layout Preservers, Page-to-File Splitting, Semantic HTML Structures, PDF Interactive Elements, Data URI Embeddings, Font Data URI Embedders.

ما هي البدائل مفتوحة المصدر لـ coolwanglu/pdf2htmlex؟

تشمل البدائل مفتوحة المصدر لـ coolwanglu/pdf2htmlex: py-pdf/pypdf — pypdf is a Python library for parsing, manipulating, and generating PDF documents. It provides high-level operations… jung-kurt/gofpdf — This is a Go-based PDF library used for the programmatic generation of PDF documents without relying on external C… pdf2htmlex/pdf2htmlex — pdf2htmlEX is a PDF to HTML converter that transforms documents into web pages while preserving the original layout,… mwilliamson/mammoth.js. microsoft/frontend-bootcamp — Frontend Workshop from HTML/CSS/JS to TypeScript/React/Redux. prawnpdf/prawn — Prawn is a Ruby library and document layout tool used for the programmatic generation of PDF files. It functions as a…

بدائل مفتوحة المصدر لـ Pdf2htmlEX

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Pdf2htmlEX.
  • py-pdf/pypdfالصورة الرمزية لـ py-pdf

    py-pdf/pypdf

    9,818عرض على GitHub↗

    pypdf is a Python library for parsing, manipulating, and generating PDF documents. It provides high-level operations for document processing, such as merging multiple files into one or splitting a single document into smaller files. The project includes specialized tools for managing interactive elements, including the creation and modification of annotations, hyperlinks, and form fields. It also supports advanced metadata management, allowing for the extraction and modification of standard document properties and XML-based XMP metadata. Beyond basic structural changes, the library covers pa

    Pythonhelp-wantedpdfpdf-documents
    عرض على GitHub↗9,818
  • jung-kurt/gofpdfالصورة الرمزية لـ jung-kurt

    jung-kurt/gofpdf

    4,468عرض على GitHub↗

    This is a Go-based PDF library used for the programmatic generation of PDF documents without relying on external C dependencies. It functions as a document generator, layout engine, security tool, and vector graphics engine for creating files containing text, images, and geometric shapes. The project distinguishes itself through a cell-based layout engine that manages automatic text wrapping, page breaks, and structured positioning. It provides specialized capabilities for vector graphics, including the rendering of Bézier curves and polygons, as well as a security toolkit for applying passwo

    Go
    عرض على GitHub↗4,468
  • pdf2htmlex/pdf2htmlexالصورة الرمزية لـ pdf2htmlEX

    pdf2htmlEX/pdf2htmlEX

    5,412عرض على GitHub↗

    pdf2htmlEX is a PDF to HTML converter that transforms documents into web pages while preserving the original layout, fonts, and formatting. It functions as a layout engine and text extractor, mapping PDF coordinate data to HTML and CSS to maintain visual fidelity. The tool converts PDF content into searchable and selectable native HTML text by embedding original document fonts. It maintains document interactivity by preserving internal links, bookmarks, and outlines, converting them into functional web navigation. The conversion process supports flexible output structures, allowing documents

    HTMLhtmlpdfpdf-document-processor
    عرض على GitHub↗5,412
  • mwilliamson/mammoth.jsالصورة الرمزية لـ mwilliamson

    mwilliamson/mammoth.js

    6,101عرض على GitHub↗
    JavaScript
    عرض على GitHub↗6,101
  • عرض جميع البدائل الـ 30 لـ Pdf2htmlEX→