awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
PDFCraftTool avatar

PDFCraftTool/pdfcraft

0
View on GitHub↗
3,113 stars·809 forks·JavaScript·agpl-3.0·36 viewspdfcraft.devtoolcafe.com↗

Pdfcraft

Pdfcraft is a containerized service for self-managed PDF processing, editing, and conversion. It provides a toolkit for document manipulation, a multi-format converter, and OCR software to transform scanned documents into searchable and editable text.

The project features a visual, node-based workflow editor that allows users to build automated pipelines by chaining together various PDF conversion and optimization operations.

The service covers a broad range of capabilities, including document management for merging and splitting files, format conversion between PDFs and office documents or images, and security tools for encryption and metadata removal. It also includes utilities for content editing, interactive form creation, and file optimization.

Features

  • Self-Hosted PDF Suites - Provides a comprehensive, self-hosted suite for managing and processing PDF files on private infrastructure.
  • Workflow Builders - Provides a drag-and-drop visual editor for building automated PDF processing pipelines.
  • PDF Page Organizers - Provides tools to combine, split, reorder, rotate, and extract pages to restructure PDF files.
  • PDF Format Converters - Converts PDF files to and from images, office documents, and structured JSON formats.
  • PDF Manipulation Utilities - Provides a wide range of utilities for merging, splitting, and restructuring PDF documents.
  • PDF Image Conversion - Extracts pages from PDF documents and saves them as JPG image files.
  • Document Exports - Provides conversion of PDF pages into editable office documents, spreadsheets, and structured JSON data.
  • PDF Workflow Orchestrators - Enables the chaining of multiple PDF operations into automated, repeatable workflows.
  • PDF Page Extraction - Divides a single PDF document into multiple smaller files based on user requirements.
  • PDF Content Editing - Enables direct modification of text and images, adding watermarks, page numbers, and electronic signatures.
  • Searchable PDF Generation - Transforms non-searchable scanned PDFs into searchable documents using an OCR-derived text layer.
  • PDF Document Management - Provides a comprehensive toolkit for merging, splitting, rotating, and reordering PDF pages.
  • PDF Document Merging - Combines multiple PDF documents into a single file using browser-side processing.
  • Containerized Application Deployments - Offers full containerization for flexible self-managed hosting across different infrastructures.
  • Containerized Deployments - Supplies a portable container image for consistent hosting of the PDF processing environment.
  • Optical Character Recognition - Uses optical character recognition to create searchable text layers from scanned PDF images.
  • Visual Node Orchestration - Features a visual graph editor to sequence PDF operations via interconnected functional nodes.
  • Content Redaction - Modifies PDF content through permanent redaction, highlighting, and adding annotations or signatures.
  • PDF Security Management - Secures PDF documents through password encryption, access permission management, and metadata removal.
  • PDF Compression - Reduces the file size of PDF documents to optimize storage and transfer speeds.
  • File Linearization - Repairs corrupted documents and linearizes files to standardize page dimensions for faster web viewing.
  • File Size Optimizations - Includes tools for compressing PDF files and optimizing them for faster web viewing.
  • Interactive PDF Design - Creates PDFs with fillable form fields, interactive text boxes, and dropdown menus.
  • Client-Side Media Processing - Provides capabilities to manipulate and convert files directly in the browser to ensure data privacy.
  • PDF - Listed in the “PDF 工具” section of the Great Open Source Project awesome list.

Star history

Star history chart for pdfcrafttool/pdfcraftStar history chart for pdfcrafttool/pdfcraft

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does pdfcrafttool/pdfcraft do?

Pdfcraft is a containerized service for self-managed PDF processing, editing, and conversion. It provides a toolkit for document manipulation, a multi-format converter, and OCR software to transform scanned documents into searchable and editable text.

What are the main features of pdfcrafttool/pdfcraft?

The main features of pdfcrafttool/pdfcraft are: Self-Hosted PDF Suites, Workflow Builders, PDF Page Organizers, PDF Format Converters, PDF Manipulation Utilities, PDF Image Conversion, Document Exports, PDF Workflow Orchestrators.

Which projects share features with pdfcrafttool/pdfcraft?

Projects with overlapping indexed features include: torakiki/pdfsam — pdfsam is a PDF manipulation software and desktop application designed for splitting, merging, rotating, and… kevin2li/pdf-guru — PDF-Guru is an AI-powered document processor and study material converter designed to transform textbooks, research… pymupdf/pymupdf — PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool,… stirling-tools/stirling-pdf — Stirling-PDF is a self-hosted document processing suite designed for secure, private file management. It functions as… pdfarranger/pdfarranger — Pdfarranger is a PDF page organizer, document editor, image converter, and booklet generator. It provides a visual… frooodle/stirling-pdf — Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents.…

Projects sharing features with Pdfcraft

These projects share indexed features with Pdfcraft. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • torakiki/pdfsamtorakiki avatar

    torakiki/pdfsam

    4,457View on GitHub↗

    pdfsam is a PDF manipulation software and desktop application designed for splitting, merging, rotating, and extracting pages from PDF documents. It functions as a PDF editor, converter, and security tool, providing capabilities to modify document structures and manage file formats. The project distinguishes itself through specialized processing capabilities, including an OCR document processor for extracting editable text from scanned images and PDF interleaving to alternate pages from multiple files. It also provides a security suite for encrypting documents, managing access permissions, an

    Javacombineextractjava
    View on GitHub↗4,457
  • kevin2li/pdf-gurukevin2li avatar

    kevin2li/PDF-Guru

    4,113View on GitHub↗

    PDF-Guru is an AI-powered document processor and study material converter designed to transform textbooks, research papers, and multimedia content into structured flashcards for spaced repetition systems like Anki. It functions as a content pipeline that uses language models to extract key concepts and facts from unstructured documents to generate question-and-answer pairs, cloze deletions, and multiple-choice cards. The system distinguishes itself through a comprehensive PDF management suite and multi-format parsing. It provides advanced document utilities including optical character recogni

    Vueai-flashcardsanki-flashcardsanki-to-pdf
    View on GitHub↗4,113
  • pymupdf/pymupdfpymupdf avatar

    pymupdf/PyMuPDF

    9,086View on GitHub↗

    PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool, OCR engine, and image converter, providing a programmatic interface to edit, merge, split, and optimize PDF and Office documents. The project distinguishes itself through high-performance capabilities, including the use of C-bindings for low-level manipulation and parallelized page processing to accelerate workloads. It provides specialized conversion paths, such as transforming PDF content into Markdown for retrieval-augmented generation and large language model pipelines. It

    Pythondata-scienceepubextract-data
    View on GitHub↗9,086
  • stirling-tools/stirling-pdfStirling-Tools avatar

    Stirling-Tools/Stirling-PDF

    81,109View on GitHub↗

    Stirling-PDF is a self-hosted document processing suite designed for secure, private file management. It functions as a comprehensive transformation engine that executes complex operations—such as merging, splitting, converting, and redacting documents—directly on the host machine. The platform provides both a browser-based interface for interactive editing and a programmatic, API-first architecture that allows for the automation of document workflows through standard HTTP requests. The project distinguishes itself through its focus on private, infrastructure-agnostic deployment and granular

    TypeScriptdockerhacktoberfestjava
    View on GitHub↗81,109
  • Compare all 30 related projects→