awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to interviewbubble/tabulo

Open-source alternatives to Tabulo

16 open-source projects similar to interviewbubble/tabulo, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Tabulo alternative.

  • frooodle/stirling-pdfAvatar de Frooodle

    Frooodle/Stirling-PDF

    81,168Voir sur GitHub↗

    Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents. It functions as a self-hosted document manager, providing a centralized interface for users to manipulate files on a private server. The system features a workflow automation engine that allows for the creation of processing pipelines to handle large volumes of documents without writing custom code. It also includes an optical character recognition tool to convert scanned PDFs into searchable and editable text. Access is managed through single sign-on integration and OIDC comp

    Java
    Voir sur GitHub↗81,168
  • guaguastandup/zotero-pdf2zhAvatar de guaguastandup

    guaguastandup/zotero-pdf2zh

    3,008Voir sur GitHub↗

    zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages using large language model APIs and machine translation services. It functions as an LLM document translator that integrates directly into Zotero libraries to process research documents. The tool distinguishes itself by generating bilingual PDFs with parallel layouts for side-by-side comparison of original and translated text. It also includes a mobile layout optimizer that crops double-column PDFs and stacks them vertically to improve readability on narrow screens. The syste

    Python
    Voir sur GitHub↗3,008
  • applicaai/digital-born-pdf-scannerA

    applicaai/digital-born-pdf-scanner

    0Voir sur GitHub↗

    Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF viewers use to display text. Some are definitely scanned documents, that do not have any text layer at all, some are searchable OCR-processed scans that contain a lot of hidden text.

    Voir sur GitHub↗0

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Find more with AI search
ckorzen/pdf-text-extraction-benchmarkC

ckorzen/pdf-text-extraction-benchmark

0Voir sur GitHub↗

This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, especially from scientific articles. It provides (1) a benchmark generator, (2) a ready-to-use benchmark and (3) an extensive evaluation, with…

Voir sur GitHub↗0
  • deepdoctection/deepdoctectionAvatar de deepdoctection

    deepdoctection/deepdoctection

    3,181Voir sur GitHub↗

    A Repo For Document AI

    Python
    Voir sur GitHub↗3,181
  • jbarlow83/ocrmypdfAvatar de jbarlow83

    jbarlow83/OCRmyPDF

    33,901Voir sur GitHub↗

    OCRmyPDF is a tool for converting image-based PDF files into machine-readable documents by adding a searchable text layer via optical character recognition. It functions as a multi-language processor capable of detecting and extracting text in over 100 different languages using linguistic data packs. The software includes a PDF image optimizer to remove image artifacts and correct page skew to improve recognition accuracy. It also provides a converter to transform scanned documents into the PDF/A standard for long-term digital archiving. The system manages PDF optimization by compressing emb

    Python
    Voir sur GitHub↗33,901
  • jorisschellekens/borbAvatar de jorisschellekens

    jorisschellekens/borb

    3,567Voir sur GitHub↗

    borb is a powerful and flexible Python library for creating and manipulating PDF files.

    Python
    Voir sur GitHub↗3,567
  • jsfenfen/parsing-prickly-pdfsAvatar de jsfenfen

    jsfenfen/parsing-prickly-pdfs

    63Voir sur GitHub↗

    Resources and worksheet for the NICAR 2016 workshop of the same name. Instructors: Jacob Fenton (jsfenfen@gmail.com) and Jeremy Singer-Vine (jsvine@gmail.com).

    Voir sur GitHub↗63
  • jsv4/opencontractsAvatar de JSv4

    JSv4/OpenContracts

    1,368Voir sur GitHub↗

    | | | | ---…

    Python
    Voir sur GitHub↗1,368
  • jsvine/pdfplumberAvatar de jsvine

    jsvine/pdfplumber

    9,732Voir sur GitHub↗

    pdfplumber is a PDF data extraction library and layout analysis tool used to retrieve text, tables, and geometric objects from PDF files using precise coordinate-based analysis. It functions as a layout analyzer and table parser that identifies the bounding boxes and visual coordinates for every character and image on a page. The library distinguishes itself through visual debugging capabilities, allowing users to render PDF pages as images and draw annotations to verify the position of extracted data. It employs line and intersection analysis to identify cell structures and convert unstructu

    Pythonpdfpdf-parsingtable-extraction
    Voir sur GitHub↗9,732
  • layout-parser/layout-parserAvatar de Layout-Parser

    Layout-Parser/layout-parser

    5,749Voir sur GitHub↗

    Layout-parser is a deep learning document layout parser and image analysis framework. It provides a toolkit for extracting structural information and layout patterns from scanned documents and digital images, transforming them into programmatic data structures for automated analysis. The framework integrates layout detection with optical character recognition to convert tabular regions into machine-readable data. It utilizes neural networks to identify and classify structural elements within document images without relying on manual rule-based systems. The system covers a broad range of docu

    Python
    Voir sur GitHub↗5,749
  • pdfminer/pdfminer.sixAvatar de pdfminer

    pdfminer/pdfminer.six

    6,906Voir sur GitHub↗

    pdfminer.six is a programmatic tool for extracting text, layout information, and metadata from PDF documents into machine-readable formats. It functions as a document parser that converts internal PDF objects and structures into accessible data objects for analysis. The project includes utilities for decrypting RC4 and AES encrypted files to enable content extraction. It also provides a layout analyzer to identify fonts, colors, and text locations to determine the organizational structure of pages. The system covers a broad range of extraction capabilities, including the retrieval of embedde

    Pythonparserpdfpython
    Voir sur GitHub↗6,906
  • uglytoad/pdfpigAvatar de UglyToad

    UglyToad/PdfPig

    2,466Voir sur GitHub↗

    Read and extract text and other content from PDFs in C# (port of PDFBox)

    C#alto-xmlcsharpdocument-analysis
    Voir sur GitHub↗2,466
  • allenai/pawlsAvatar de allenai

    allenai/pawls

    431Voir sur GitHub↗

    Demo Server | Video Tutorial | Paper

    Python
    Voir sur GitHub↗431
  • xyntopia/pydoxtoolsX

    xyntopia/pydoxtools

    0Voir sur GitHub↗

    title: 'pydoxtools (Python Library)' library name: pydoxtools keywords: pydoxtools, AI, AI-Composition, ETL, pipelines, knowledge graphs

    Voir sur GitHub↗0
  • apache/pdfboxAvatar de apache

    apache/pdfbox

    3,079Voir sur GitHub↗

    Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file distributed with this work for additional information regarding copyright ownership. The ASF licenses this file to You under the Apache License, Version 2.0 (the "License"); you…

    Java
    Voir sur GitHub↗3,079