awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
the-paperless-project avatar

the-paperless-project/paperlessArchived

0
View on GitHub↗
7,917 स्टार्स·500 फोर्क्स·Python·GPL-3.0·8 व्यूज़

Paperless

Paperless is a self-hosted document management system designed to digitize, index, and archive paper documents. It functions as an optical character recognition system that converts scanned images and PDFs into a searchable digital library, providing a web-based interface for querying and retrieving documents from a database.

The system features an automated file ingestion pipeline that monitors specific directories and email inboxes to process and import documents without manual uploading. To maintain a private archive, it includes on-disk encryption for sensitive files and the ability to organize physical storage using metadata-driven filename templates.

The platform covers broad capabilities for document processing, including image cleaning to remove speckles and correct skewing for better text recognition. It also provides tools for exporting archived documents to local directories for external backups and allows for user interface customization via custom styles and scripts.

The application is packaged as a containerized deployment to ensure consistent installation across different environments.

Features

  • Document Digitization Tools - Converts physical papers into a searchable digital library through scanning and indexing for long-term storage.
  • Optical Character Recognition - Converts scanned image-based documents into searchable text using optical character recognition.
  • Optical Character Recognition - Extracts searchable text from scanned images and PDFs using optical character recognition.
  • Email-Based Ingestion - Monitors a specific email inbox to automatically ingest documents containing a designated secret key.
  • Document Archiving Systems - Digitizes, indexes, and archives paper documents for long-term storage and efficient retrieval.
  • File Ingestion Services - Automatically ingests documents from monitored directories and email inboxes for processing and archiving.
  • Automated Document Ingestion - Monitors specific directories for new files to automatically ingest and process them.
  • Document Ingestion Pipelines - Ships a workflow that monitors directories and email to automatically import and process scanned documents.
  • Full Text Search - Indexes extracted text and document metadata in a database to enable rapid full-text retrieval.
  • Optical Character Recognition - Extracts text from scanned images via OCR to enable high-speed full-text search.
  • Search Indexing - Provides a web interface to query the indexed database and retrieve specific digitized documents.
  • Document Management Systems - Functions as a complete system for digitizing, indexing, and archiving paper documents using OCR.
  • Document Managers - Provides a private, self-hosted system to organize, tag, and store digital paper records.
  • Image Pre-processing - Removes speckles and corrects skewing in scanned images to improve text recognition accuracy.
  • Metadata Template Resolvers - Generates physical filenames from document metadata using customizable templates for organized storage.
  • Secure Storage - Protects sensitive digitized documents using on-disk encryption and secure network access.
  • Document Encryption - Implements on-disk encryption with on-the-fly decryption during document downloads.
  • Storage Encryption - Provides on-disk encryption for sensitive documents and performs decryption during user downloads.
  • Data Management Systems - System for indexing and archiving scanned paper documents.
  • File Sharing Tools - Archived tool for scanning and indexing documents.
  • डॉक्यूमेंटेशन और नॉलेज - System for indexing and archiving scanned documents.

स्टार हिस्ट्री

the-paperless-project/paperless के लिए स्टार हिस्ट्री चार्टthe-paperless-project/paperless के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Paperless के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Paperless के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • awesome-selfhosted/awesome-selfhostedawesome-selfhosted का अवतार

    awesome-selfhosted/awesome-selfhosted

    299,516GitHub पर देखें↗

    This project is a community-curated directory of open-source software designed for deployment in private server environments and home labs. It serves as a comprehensive resource for discovering independent, self-hosted alternatives to mainstream cloud services, enabling users to maintain full data ownership and control over their digital infrastructure. The directory is structured through a hierarchical taxonomy that organizes a vast collection of applications into logical categories, ranging from media management and data analytics to private communication and team productivity tools. It dis

    awesomeawesome-listcloud
    GitHub पर देखें↗299,516
  • papra-hq/paprapapra-hq का अवतार

    papra-hq/papra

    3,838GitHub पर देखें↗

    Papra is a self-hosted document management system designed for digital archiving, organization, and retrieval. It serves as a centralized platform for storing files with a focus on security, providing an encrypted file archive using AES-256-GCM and a programmatic interface for managing documents and metadata via a REST API, SDK, and command line tools. The system distinguishes itself through an automated document ingestion engine that imports files via email forwarding, monitored folders, and webhook listeners. It further enhances discoverability by acting as an OCR document indexer, extracti

    TypeScriptapparchivedocument
    GitHub पर देखें↗3,838
  • miniflux/v2miniflux का अवतार

    miniflux/v2

    9,389GitHub पर देखें↗

    This project is a self-hosted RSS feed aggregator and reader designed to collect and organize content from RSS, Atom, and JSON feeds. It functions as a privacy-focused client that blocks pixel trackers and strips URL parameters to prevent third-party tracking and referrer leakage. The system is built as a REST API feed reader, exposing its data and user accounts through a programmable interface for third-party clients. It maintains compatibility with the OPML standard for importing and exporting subscriptions and provides tools for web content extraction using readability parsers and custom r

    Goatomfeedgo
    GitHub पर देखें↗9,389
  • gollum/gollumgollum का अवतार

    gollum/gollum

    14,279GitHub पर देखें↗

    Gollum is a Git-powered wiki engine and content management system that provides a web-based interface for editing and organizing files stored in a Git repository. It functions as a self-hosted documentation tool, using a Git-based storage backend to manage page content and track version history. The system is characterized by a pluggable markup rendering architecture that converts multiple markup languages and specialized notations into HTML. It supports a wide array of rich content, including mathematical typesetting, BibTeX bibliographies, and diagrams rendered via Mermaid. Broad capabilit

    Rubydocumentationdocumentation-toolgollum
    GitHub पर देखें↗14,279
Paperless के सभी 30 विकल्प देखें→

अक्सर पूछे जाने वाले प्रश्न

the-paperless-project/paperless क्या करता है?

Paperless is a self-hosted document management system designed to digitize, index, and archive paper documents. It functions as an optical character recognition system that converts scanned images and PDFs into a searchable digital library, providing a web-based interface for querying and retrieving documents from a database.

the-paperless-project/paperless की मुख्य विशेषताएं क्या हैं?

the-paperless-project/paperless की मुख्य विशेषताएं हैं: Document Digitization Tools, Optical Character Recognition, Email-Based Ingestion, Document Archiving Systems, File Ingestion Services, Automated Document Ingestion, Document Ingestion Pipelines, Full Text Search।

the-paperless-project/paperless के कुछ ओपन-सोर्स विकल्प क्या हैं?

the-paperless-project/paperless के ओपन-सोर्स विकल्पों में शामिल हैं: awesome-selfhosted/awesome-selfhosted — This project is a community-curated directory of open-source software designed for deployment in private server… papra-hq/papra — Papra is a self-hosted document management system designed for digital archiving, organization, and retrieval. It… miniflux/v2 — This project is a self-hosted RSS feed aggregator and reader designed to collect and organize content from RSS, Atom,… gollum/gollum — Gollum is a Git-powered wiki engine and content management system that provides a web-based interface for editing and… tesseract-ocr/tessdata — This repository provides the pre-trained neural network and legacy data files used by Tesseract to recognize and… pymupdf/pymupdf — PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool,…