awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Python PDF Libraries

Clasament actualizat la 13 iul. 2026

For a python library for manipulating pdf files, the strongest matches are mstamy2/pypdf2 (This library provides a robust set of tools for), py-pdf/pypdf2 (This library provides a robust set of tools for) and pymupdf/pymupdf (PyMuPDF is a high-performance Python library that provides a). jsvine/pdfplumber and jbarlow83/ocrmypdf round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Selectăm repository-uri open-source de pe GitHub care se potrivesc cu „best python pdf libraries”. Rezultatele sunt clasificate după relevanța față de căutarea ta — folosește filtrele de mai jos pentru a rafina rezultatele sau utilizează AI-ul.

Python PDF Libraries

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • mstamy2/pypdf2Avatar mstamy2

    mstamy2/PyPDF2

    10,064Vezi pe GitHub↗

    PyPDF2 is a pure Python library for reading, writing, and manipulating PDF files. It functions as a document manipulator, text extractor, and encryption tool, allowing users to process PDF files without relying on external C libraries or native binaries. The library provides specialized tools for modifying document structures, such as merging multiple files into one, splitting documents into separate files, and transforming page layouts through cropping. It also includes capabilities for securing documents via passwords and encryption. Additional capabilities include the extraction of writte

    This library provides a robust set of tools for reading, writing, merging, and splitting PDF documents, making it a core utility for programmatic PDF manipulation and text extraction in Python.

    PythonPDF LibrariesPDF Manipulation UtilitiesText Extraction
    Vezi pe GitHub↗10,064
  • py-pdf/pypdf2Avatar py-pdf

    py-pdf/PyPDF2

    10,094Vezi pe GitHub↗

    PyPDF2 is a pure Python library for transforming, securing, and extracting data from PDF documents. It provides a comprehensive suite of tools to modify page layouts, manage document security, and retrieve embedded metadata without relying on external C libraries. The toolkit enables document assembly through the merging of multiple files and the splitting of documents into smaller parts. It also supports page-level transformations, including the ability to rotate pages and adjust visible crop areas. The library includes capabilities for security management via password-based encryption and

    This library provides a robust set of tools for PDF manipulation, merging, splitting, and text extraction, making it a core utility for programmatic PDF processing in Python.

    PythonPDF LibrariesPDF Manipulation LibrariesPDF Parsers
    Vezi pe GitHub↗10,094
  • pymupdf/pymupdfAvatar pymupdf

    pymupdf/PyMuPDF

    9,086Vezi pe GitHub↗

    PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool, OCR engine, and image converter, providing a programmatic interface to edit, merge, split, and optimize PDF and Office documents. The project distinguishes itself through high-performance capabilities, including the use of C-bindings for low-level manipulation and parallelized page processing to accelerate workloads. It provides specialized conversion paths, such as transforming PDF content into Markdown for retrieval-augmented generation and large language model pipelines. It

    PyMuPDF is a high-performance Python library that provides a comprehensive suite of tools for PDF generation, text extraction, manipulation, and OCR integration, making it a flagship solution for programmatic document processing.

    PythonOptical Character RecognitionOptical Character RecognitionPDF Form Filling
    Vezi pe GitHub↗9,086
  • jsvine/pdfplumberAvatar jsvine

    jsvine/pdfplumber

    9,732Vezi pe GitHub↗

    pdfplumber is a PDF data extraction library and layout analysis tool used to retrieve text, tables, and geometric objects from PDF files using precise coordinate-based analysis. It functions as a layout analyzer and table parser that identifies the bounding boxes and visual coordinates for every character and image on a page. The library distinguishes itself through visual debugging capabilities, allowing users to render PDF pages as images and draw annotations to verify the position of extracted data. It employs line and intersection analysis to identify cell structures and convert unstructu

    This library is a specialized tool for extracting text, tables, and geometric data from PDFs, fitting the category well despite its primary focus on extraction rather than document generation or form filling.

    PythonPDF LibrariesPDF ParsersText Extraction
    Vezi pe GitHub↗9,732
  • jbarlow83/ocrmypdfAvatar jbarlow83

    jbarlow83/OCRmyPDF

    33,901Vezi pe GitHub↗

    OCRmyPDF is a tool for converting image-based PDF files into machine-readable documents by adding a searchable text layer via optical character recognition. It functions as a multi-language processor capable of detecting and extracting text in over 100 different languages using linguistic data packs. The software includes a PDF image optimizer to remove image artifacts and correct page skew to improve recognition accuracy. It also provides a converter to transform scanned documents into the PDF/A standard for long-term digital archiving. The system manages PDF optimization by compressing emb

    This tool is a specialized Python-based utility for OCR-based text layer addition and PDF/A conversion, making it a highly effective library for specific PDF manipulation and archival tasks.

    PythonOptical Character Recognition EnginesPDF Generation
    Vezi pe GitHub↗33,901
  • vikparuchuri/markerAvatar VikParuchuri

    VikParuchuri/marker

    36,164Vezi pe GitHub↗

    Marker is an LLM-powered document parser and OCR pipeline designed to convert PDFs and unstructured files into structured markdown, JSON, and HTML. It functions as a data preprocessor that transforms complex documents into machine-readable formats while preserving tables, equations, and layout structures. The system utilizes large language models to refine OCR accuracy, clean mathematical notation, and merge fragmented tables across multiple pages. It employs model-based layout analysis to predict block types and bounding boxes, ensuring a more precise conversion of document elements. Capabi

    Marker is a specialized Python library for extracting structured data and text from PDFs using LLM-powered OCR and layout analysis, making it a highly effective tool for the extraction and conversion aspects of PDF processing.

    PythonOptical Character RecognitionOptical Character Recognition Engines
    Vezi pe GitHub↗36,164
  • pdfminer/pdfminer.sixAvatar pdfminer

    pdfminer/pdfminer.six

    6,906Vezi pe GitHub↗

    pdfminer.six is a programmatic tool for extracting text, layout information, and metadata from PDF documents into machine-readable formats. It functions as a document parser that converts internal PDF objects and structures into accessible data objects for analysis. The project includes utilities for decrypting RC4 and AES encrypted files to enable content extraction. It also provides a layout analyzer to identify fonts, colors, and text locations to determine the organizational structure of pages. The system covers a broad range of extraction capabilities, including the retrieval of embedde

    This library is a specialized tool for parsing and extracting text, layout, and metadata from PDF documents, though it focuses on analysis and extraction rather than PDF generation or creation.

    PythonPDF Parsers
    Vezi pe GitHub↗6,906
  • py-pdf/pypdfAvatar py-pdf

    py-pdf/pypdf

    9,818Vezi pe GitHub↗

    pypdf is a Python library for parsing, manipulating, and generating PDF documents. It provides high-level operations for document processing, such as merging multiple files into one or splitting a single document into smaller files. The project includes specialized tools for managing interactive elements, including the creation and modification of annotations, hyperlinks, and form fields. It also supports advanced metadata management, allowing for the extraction and modification of standard document properties and XML-based XMP metadata. Beyond basic structural changes, the library covers pa

    This library is a comprehensive tool for programmatic PDF generation, manipulation, and text extraction, though it lacks built-in OCR capabilities.

    PythonPDF LibrariesPDF Manipulation UtilitiesText Extraction
    Vezi pe GitHub↗9,818
  • camelot-dev/camelotAvatar camelot-dev

    camelot-dev/camelot

    3,764Vezi pe GitHub↗

    Camelot is a Python library and processing engine designed to extract tabular data from PDF documents. It converts unstructured tables into machine-readable formats such as CSV, JSON, and Excel. The project provides specialized toolsets for different document types, using line detection for ruled tables and whitespace analysis for borderless tables. It includes an optical character recognition system to recover structured data from image-based scanned PDFs that lack a digital text layer. The library handles complex document layouts, including encrypted files, rotated pages, and tables that s

    Camelot is a specialized Python library focused on the extraction of tabular data from PDFs, which fits the category despite its narrow scope of focusing on tables rather than general-purpose PDF generation or manipulation.

    PythonPDF Parsers
    Vezi pe GitHub↗3,764
  • opendatalab/mineruAvatar opendatalab

    opendatalab/MinerU

    67,734Vezi pe GitHub↗

    MinerU is a document parsing pipeline designed to transform unstructured files into machine-readable, structured data. It utilizes deep learning models to perform layout analysis, identifying document regions and extracting complex content such as mathematical expressions. By combining these neural network inferences with geometric heuristics, the system reconstructs the reading order and structural hierarchy of documents to ensure accurate data representation. The project distinguishes itself through a multi-stage processing workflow that integrates layout detection, optical character recogn

    MinerU is a specialized Python-based document parsing pipeline that excels at extracting structured data and text from complex PDFs using deep learning, making it a powerful tool for PDF data extraction and analysis.

    PythonDeployment & ServingDocument Layout AnalysisAutomated Data Extraction
    Vezi pe GitHub↗67,734
  • docling-project/doclingAvatar docling-project

    docling-project/docling

    61,674Vezi pe GitHub↗

    Docling is a modular framework designed for document parsing, layout analysis, and structured data extraction. It transforms unstructured files and web content into a unified, hierarchical data model that preserves the spatial and semantic relationships between text, tables, images, and layout elements. By normalizing diverse input formats into a consistent internal representation, the library enables uniform processing across various document types. The project distinguishes itself through a schema-driven approach that maps document regions to strongly-typed objects, ensuring data accuracy t

    Docling is a powerful Python framework for parsing and extracting structured data from PDFs, making it a highly capable tool for document processing even though its primary focus is on intelligent extraction rather than PDF generation or form filling.

    PythonDocument and LLM PreparationDocument Layout AnalyzersHierarchical Document Models
    Vezi pe GitHub↗61,674
  • pikepdf/pikepdfAvatar pikepdf

    pikepdf/pikepdf

    2,744Vezi pe GitHub↗

    A Python library for reading and writing PDF, powered by QPDF

    This library provides robust capabilities for reading, writing, and manipulating PDF files by leveraging the QPDF engine, making it a strong choice for programmatic PDF processing despite lacking built-in OCR or form-filling features.

    PythonDocument and File Processing
    Vezi pe GitHub↗2,744
Compară top 10 dintr-o privire
RepositorySteleLimbajLicențăUltimul push
mstamy2/pypdf210.1KPythonNOASSERTION16 iun. 2026
py-pdf/pypdf210.1KPythonNOASSERTION25 iun. 2026
pymupdf/pymupdf9.1KPythonagpl-3.018 feb. 2026
jsvine/pdfplumber9.7KPythonmit28 ian. 2026
jbarlow83/ocrmypdf33.9KPythonMPL-2.017 iun. 2026
vikparuchuri/marker36.2KPythonGPL-3.06 iun. 2026
pdfminer/pdfminer.six6.9KPythonmit13 feb. 2026
py-pdf/pypdf9.8KPythonother19 feb. 2026
camelot-dev/camelot3.8KPythonMIT24 iun. 2026
opendatalab/mineru67.7KPythonNOASSERTION15 iun. 2026

Related searches

  • a python library for manipulating pdf files
  • an open source tool for editing PDFs
  • a library for rendering PDFs on Android
  • a javascript library for generating pdf files
  • a library for generating pdfs in C#
  • librărie de generare PDF pentru aplicații Java
  • instrument open source pentru editarea fișierelor PDF
  • editor open source pentru documente PDF