awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to interviewbubble/tabulo

Open-source alternatives to Tabulo

16 open-source projects similar to interviewbubble/tabulo, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Tabulo alternative.

  • frooodle/stirling-pdfAvatar de Frooodle

    Frooodle/Stirling-PDF

    81,168Ver en GitHub↗

    Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents. It functions as a self-hosted document manager, providing a centralized interface for users to manipulate files on a private server. The system features a workflow automation engine that allows for the creation of processing pipelines to handle large volumes of documents without writing custom code. It also includes an optical character recognition tool to convert scanned PDFs into searchable and editable text. Access is managed through single sign-on integration and OIDC comp

    Java
    Ver en GitHub↗81,168
  • guaguastandup/zotero-pdf2zhAvatar de guaguastandup

    guaguastandup/zotero-pdf2zh

    3,008Ver en GitHub↗

    zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages using large language model APIs and machine translation services. It functions as an LLM document translator that integrates directly into Zotero libraries to process research documents. The tool distinguishes itself by generating bilingual PDFs with parallel layouts for side-by-side comparison of original and translated text. It also includes a mobile layout optimizer that crops double-column PDFs and stacks them vertically to improve readability on narrow screens. The syste

    Python
    Ver en GitHub↗3,008
  • applicaai/digital-born-pdf-scannerA

    applicaai/digital-born-pdf-scanner

    0Ver en GitHub↗

    Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF viewers use to display text. Some are definitely scanned documents, that do not have any text layer at all, some are searchable OCR-processed scans that contain a lot of hidden text.

    Ver en GitHub↗0

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Find more with AI search
ckorzen/pdf-text-extraction-benchmarkC

ckorzen/pdf-text-extraction-benchmark

0Ver en GitHub↗

This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, especially from scientific articles. It provides (1) a benchmark generator, (2) a ready-to-use benchmark and (3) an extensive evaluation, with…

Ver en GitHub↗0
  • deepdoctection/deepdoctectionAvatar de deepdoctection

    deepdoctection/deepdoctection

    3,181Ver en GitHub↗

    A Repo For Document AI

    Python
    Ver en GitHub↗3,181
  • jbarlow83/ocrmypdfAvatar de jbarlow83

    jbarlow83/OCRmyPDF

    33,901Ver en GitHub↗

    OCRmyPDF is a tool for converting image-based PDF files into machine-readable documents by adding a searchable text layer via optical character recognition. It functions as a multi-language processor capable of detecting and extracting text in over 100 different languages using linguistic data packs. The software includes a PDF image optimizer to remove image artifacts and correct page skew to improve recognition accuracy. It also provides a converter to transform scanned documents into the PDF/A standard for long-term digital archiving. The system manages PDF optimization by compressing emb

    Python
    Ver en GitHub↗33,901
  • jorisschellekens/borbAvatar de jorisschellekens

    jorisschellekens/borb

    3,567Ver en GitHub↗

    borb is a powerful and flexible Python library for creating and manipulating PDF files.

    Python
    Ver en GitHub↗3,567
  • jsfenfen/parsing-prickly-pdfsAvatar de jsfenfen

    jsfenfen/parsing-prickly-pdfs

    63Ver en GitHub↗

    Resources and worksheet for the NICAR 2016 workshop of the same name. Instructors: Jacob Fenton (jsfenfen@gmail.com) and Jeremy Singer-Vine (jsvine@gmail.com).

    Ver en GitHub↗63
  • jsv4/opencontractsAvatar de JSv4

    JSv4/OpenContracts

    1,368Ver en GitHub↗

    | | | | ---…

    Python
    Ver en GitHub↗1,368
  • jsvine/pdfplumberAvatar de jsvine

    jsvine/pdfplumber

    9,732Ver en GitHub↗

    pdfplumber is a PDF data extraction library and layout analysis tool used to retrieve text, tables, and geometric objects from PDF files using precise coordinate-based analysis. It functions as a layout analyzer and table parser that identifies the bounding boxes and visual coordinates for every character and image on a page. The library distinguishes itself through visual debugging capabilities, allowing users to render PDF pages as images and draw annotations to verify the position of extracted data. It employs line and intersection analysis to identify cell structures and convert unstructu

    Pythonpdfpdf-parsingtable-extraction
    Ver en GitHub↗9,732
  • layout-parser/layout-parserAvatar de Layout-Parser

    Layout-Parser/layout-parser

    5,749Ver en GitHub↗

    Layout-parser is a deep learning document layout parser and image analysis framework. It provides a toolkit for extracting structural information and layout patterns from scanned documents and digital images, transforming them into programmatic data structures for automated analysis. The framework integrates layout detection with optical character recognition to convert tabular regions into machine-readable data. It utilizes neural networks to identify and classify structural elements within document images without relying on manual rule-based systems. The system covers a broad range of docu

    Python
    Ver en GitHub↗5,749
  • pdfminer/pdfminer.sixAvatar de pdfminer

    pdfminer/pdfminer.six

    6,906Ver en GitHub↗

    pdfminer.six is a programmatic tool for extracting text, layout information, and metadata from PDF documents into machine-readable formats. It functions as a document parser that converts internal PDF objects and structures into accessible data objects for analysis. The project includes utilities for decrypting RC4 and AES encrypted files to enable content extraction. It also provides a layout analyzer to identify fonts, colors, and text locations to determine the organizational structure of pages. The system covers a broad range of extraction capabilities, including the retrieval of embedde

    Pythonparserpdfpython
    Ver en GitHub↗6,906
  • uglytoad/pdfpigAvatar de UglyToad

    UglyToad/PdfPig

    2,466Ver en GitHub↗

    Read and extract text and other content from PDFs in C# (port of PDFBox)

    C#alto-xmlcsharpdocument-analysis
    Ver en GitHub↗2,466
  • allenai/pawlsAvatar de allenai

    allenai/pawls

    431Ver en GitHub↗

    Demo Server | Video Tutorial | Paper

    Python
    Ver en GitHub↗431
  • xyntopia/pydoxtoolsX

    xyntopia/pydoxtools

    0Ver en GitHub↗

    title: 'pydoxtools (Python Library)' library name: pydoxtools keywords: pydoxtools, AI, AI-Composition, ETL, pipelines, knowledge graphs

    Ver en GitHub↗0
  • apache/pdfboxAvatar de apache

    apache/pdfbox

    3,079Ver en GitHub↗

    Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file distributed with this work for additional information regarding copyright ownership. The ASF licenses this file to You under the Apache License, Version 2.0 (the "License"); you…

    Java
    Ver en GitHub↗3,079