awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
grobidOrg avatar

grobidOrg/grobid

0
View on GitHub↗
4,954 स्टार्स·559 फोर्क्स·Java·Apache-2.0·7 व्यूज़grobid.readthedocs.io↗

Grobid

Grobid एक मशीन लर्निंग सिस्टम है जिसे शैक्षणिक और वैज्ञानिक PDF प्रकाशनों को संरचित XML में बदलने के लिए डिज़ाइन किया गया है। यह एक PDF से XML पार्सर और स्कॉलरली मेटाडेटा एक्सट्रैक्टर के रूप में कार्य करता है, जो शोध पत्रों से शीर्षकों, लेखकों, संबद्धताओं और ग्रंथ सूची संबंधी संदर्भों की पहचान और सामान्यीकरण करता है।

सिस्टम कच्चे PDF को कार्यात्मक क्षेत्रों में विभाजित करने के लिए एक डीप लर्निंग डॉक्यूमेंट सेगमेंट का उपयोग करता है और मेटाडेटा संवर्धन और DOI रिज़ॉल्यूशन के लिए बाहरी रजिस्ट्रियों के खिलाफ उद्धरणों का मिलान करने के लिए एक ग्रंथ सूची संदर्भ रिज़ॉल्वर को नियोजित करता है। यह एक पूर्ण मशीन लर्निंग मॉडल ट्रेनिंग पाइपलाइन का समर्थन करता है, जो एनोटेटेड ट्रेनिंग कॉर्पोरा के निर्माण, मॉडल रिट्रेनिंग और मॉडल बाइनरीज़ के निर्यात की अनुमति देता है।

यह प्रोजेक्ट डॉक्यूमेंट हेडर पार्सिंग, पूर्ण-टेक्स्ट बॉडी स्ट्रक्चरिंग और फंडिंग जानकारी तथा पेटेंट उद्धरणों जैसे डोमेन-विशिष्ट संस्थाओं की पहचान सहित निष्कर्षण क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह बाउंडिंग बॉक्स निष्कर्षण और मूल PDF लेआउट के साथ सिमेंटिक लेबल्स को सिंक्रोनाइज़ करने के लिए कोऑर्डिनेट मैपिंग के लिए स्थानिक विश्लेषण उपकरण भी प्रदान करता है।

एप्लिकेशन को कंटेनराइज़्ड इमेजेस के माध्यम से डिप्लॉय किया जा सकता है और इसमें बड़े डॉक्यूमेंट कलेक्शन्स की मल्टी-थ्रेडेड बैच प्रोसेसिंग के लिए कमांड-लाइन यूटिलिटीज शामिल हैं।

Features

  • PDF Semantic Parsers - Provides a machine learning system designed to extract semantic entities and technical metadata from PDF files.
  • PDF to XML Converters - Transforms complete academic PDF documents, including headers, bodies, and bibliographies, into structured XML using semantic segmentation.
  • Bibliographic Metadata Extraction - Identifies and normalizes authors, affiliations, and titles from scientific documents into structured XML or BibTeX.
  • Academic Paper Metadata Extractors - Identifies and normalizes key scholarly metadata including titles, authors, and affiliations from research papers.
  • Affiliation Parsers - Converts raw organizational address strings and academic affiliations into normalized XML formats.
  • Citation Parsers - Transforms individual raw bibliographical reference strings into structured XML or BibTeX records.
  • Citation Reference Extraction - Identifies and extracts bibliographic citation fragments from the reference sections of academic documents.
  • Document Structure Analysis - Uses sequence labeling models to segment PDF documents into functional regions like headers, footers, and body text.
  • Scientific Document Segmentation - Uses deep learning to segment academic PDFs into functional regions like abstracts and bodies.
  • Document Parsing Model Training - Provides a training pipeline for neural networks designed to recognize academic document layouts and text.
  • Document Layout Analysis - Analyzes font styles and spatial positioning to identify structural hierarchies, text, and tables within complex documents.
  • Neural Bibliographic Extractors - Uses deep learning models to parse citations, affiliations, and publication headers with high accuracy.
  • Scholarly Document Digitization - Digitizes academic papers by transforming research PDFs into machine-readable formats with semantic labeling and mapped coordinates.
  • Bibliographic Reference Parsing - Deconstructs raw bibliographic reference strings from PDFs into structured metadata and resolves them via APIs.
  • Citation Analysis Workflows - Links inline citation callouts to full bibliographic references and validates them against external registries and DOIs.
  • Citation Context Resolution - Associates inline citation callouts within the text with their corresponding entries in the bibliography.
  • Document Reference Mapping - Links inline callouts and citations to their corresponding bibliographic references, figures, and tables.
  • Bibliographic Registry Resolution - Resolves extracted bibliographic data against external academic registries to normalize records and append DOIs.
  • Publication Context Extraction - Identifies journal titles, publishers, and document types to categorize the source of a publication.
  • Bibliographic Identifier Extractions - Detects and classifies strong academic identifiers such as DOI, PII, ISSN, and ISBN.
  • Document Header Extraction - Extracts titles, authors, and abstracts from the initial pages of academic PDF documents into structured formats.
  • Scholarly Entity Parsing - Recognizes and normalizes scholarly metadata strings, including physical quantities and funding information.
  • Scholarly Document Structure Parsing - Converts academic PDF publications into structured XML by extracting headers, body text, and bibliographic sections.
  • Semantic Entity Coordinate Mapping - Maps extracted semantic entities to their original PDF bounding boxes for precise visual highlighting.
  • Document Layout Bounding Box Extractors - Extracts bounding box coordinates and font styles to improve the accuracy of structural document recognition.
  • Academic Author Name Parsing - Normalizes raw strings of author and editor names from scientific headers and references into XML.
  • Body Text Segmentation - Segments the PDF body into structured elements including paragraphs, section titles, footnotes, and figures.
  • Bibliographic Metadata Enrichment - Enriches bibliographic records by matching extracted data against external services to append missing identifiers.
  • Document Pre-annotation Generators - Produces XML files from raw PDFs to serve as starting points for manual correction in the training pipeline.
  • Document Language Detection - Automatically identifies the natural language used within a document to facilitate linguistic analysis.
  • Domain-Specific Parsing Adaptations - Applies specialized machine learning processing logic tailored for specific document types like patents or medical publications.
  • Machine Learning Training - Fine-tunes document processing models using custom datasets and integrated evaluation pipelines.
  • GPU Training Accelerators - Utilizes GPU hardware to accelerate the deep learning process and reduce model training time.
  • Annotation Validation Tools - Provides interfaces for correcting and validating pre-annotated XML data side-by-side with original source PDFs.
  • Training Corpus Generators - Generates pre-annotated XML files from raw PDFs to create ground-truth training corpora for parsing models.
  • Sequence Labeling Architectures - Implements sequence labeling architectures to segment structural regions of academic documents.
  • Training Data Generators - Creates structured training data files from PDF documents to bootstrap machine learning parsing models.
  • Funding - Extracts grant and funder details from research documents and matches them against established registries.
  • Table Bounding Box Detections - Detects the overall position and boundaries of tables within document layouts using bounding box coordinates.
  • Technical Element Isolation - Isolates technical components such as mathematical formulas, figures, and list items from the main body text.
  • Publication License Identification - Detects copyright owners and specific licensing terms associated with scholarly publications.
  • Scientific Domain - Identifies specialized mentions such as software, datasets, and astronomical entities within scientific literature.
  • Administrative Metadata Extraction - Captures administrative and legal metadata including funding disclosures, copyright statements, and submission dates.
  • Patent Data Extraction - Extracts and parses structured bibliographic and reference data specifically from patent publications.
  • Parsing Model Trainers - Provides a dedicated environment for training, evaluating, and exporting custom sequence labeling models for document parsing.
  • Model Architecture Configurations - Provides a configuration system to select specific deep learning or CRF architectures for different parsing tasks.
  • Batch Document Processing - Processes large volumes of PDF files simultaneously using multi-threaded clients to increase throughput.
  • Multi-Threaded Batch Processing - Implements multi-threaded batch processing to maximize hardware throughput when parsing large document collections.
  • Hybrid Parsing Engines - Uses a hybrid parsing engine that switches between deep learning and feature-based models to balance speed and accuracy.
  • Extraction Accuracy Evaluators - Measures the precision, recall, and F1-scores of parsed data against ground-truth documents to evaluate extraction accuracy.
  • Parsing Pipeline Evaluators - Measures the end-to-end accuracy of the PDF-to-XML processing pipeline using independent holdout datasets.

स्टार हिस्ट्री

grobidorg/grobid के लिए स्टार हिस्ट्री चार्टgrobidorg/grobid के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Grobid के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Grobid के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • jabref/jabrefJabRef का अवतार

    JabRef/jabref

    4,373GitHub पर देखें↗

    This project is a desktop-based bibliographic reference manager designed to organize academic research libraries and automate citation workflows. It functions as a research assistant that integrates directly with word processors and text editors, enabling users to insert and format references while writing. The application is built on a Java-based portable runtime, allowing it to operate as a self-contained tool that stores preferences and data in local configuration files. The platform distinguishes itself through a modular plugin architecture and a commitment to human-readable, text-based f

    Javaacademiaacademic-publicationsai
    GitHub पर देखें↗4,373
  • layout-parser/layout-parserLayout-Parser का अवतार

    Layout-Parser/layout-parser

    5,749GitHub पर देखें↗

    Layout-parser is a deep learning document layout parser and image analysis framework. It provides a toolkit for extracting structural information and layout patterns from scanned documents and digital images, transforming them into programmatic data structures for automated analysis. The framework integrates layout detection with optical character recognition to convert tabular regions into machine-readable data. It utilizes neural networks to identify and classify structural elements within document images without relying on manual rule-based systems. The system covers a broad range of docu

    Python
    GitHub पर देखें↗5,749
  • funstory-ai/babeldocfunstory-ai का अवतार

    funstory-ai/BabelDOC

    7,752GitHub पर देखें↗

    BabelDOC is a technical document translation system designed to translate PDF files while preserving their original layout and styling. It functions as a layout-preserving translator that utilizes large language models to convert content into target languages, specifically tailored for scientific and technical documents. The system distinguishes itself through specialized handling of academic content, including the identification and preservation of mathematical formulas and complex layout structures. It ensures technical accuracy by employing glossary-driven terminology enforcement, using so

    Python
    GitHub पर देखें↗7,752
  • oomol-lab/pdf-craftoomol-lab का अवतार

    oomol-lab/pdf-craft

    4,867GitHub पर देखें↗

    pdf-craft is an OCR-based document parser and structure extractor designed to convert PDF files into structured data, Markdown, or EPUB ebooks. It utilizes optical character recognition and statistical analysis to identify document hierarchies and extract text and structured content. The system features specialized rendering for mathematical formulas and tables, using heuristic reconstruction to convert tabular data into digital formats. It includes a document structure extractor that builds tables of contents by analyzing font sizes, linguistic patterns, and language model title detection.

    Pythondeepseek-ocrdocumentocr
    GitHub पर देखें↗4,867
Grobid के सभी 30 विकल्प देखें→

अक्सर पूछे जाने वाले प्रश्न

grobidorg/grobid क्या करता है?

Grobid एक मशीन लर्निंग सिस्टम है जिसे शैक्षणिक और वैज्ञानिक PDF प्रकाशनों को संरचित XML में बदलने के लिए डिज़ाइन किया गया है। यह एक PDF से XML पार्सर और स्कॉलरली मेटाडेटा एक्सट्रैक्टर के रूप में कार्य करता है, जो शोध पत्रों से शीर्षकों, लेखकों, संबद्धताओं और ग्रंथ सूची संबंधी संदर्भों की पहचान और सामान्यीकरण करता है।

grobidorg/grobid की मुख्य विशेषताएं क्या हैं?

grobidorg/grobid की मुख्य विशेषताएं हैं: PDF Semantic Parsers, PDF to XML Converters, Bibliographic Metadata Extraction, Academic Paper Metadata Extractors, Affiliation Parsers, Citation Parsers, Citation Reference Extraction, Document Structure Analysis।

grobidorg/grobid के कुछ ओपन-सोर्स विकल्प क्या हैं?

grobidorg/grobid के ओपन-सोर्स विकल्पों में शामिल हैं: jabref/jabref — This project is a desktop-based bibliographic reference manager designed to organize academic research libraries and… layout-parser/layout-parser — Layout-parser is a deep learning document layout parser and image analysis framework. It provides a toolkit for… oomol-lab/pdf-craft — pdf-craft is an OCR-based document parser and structure extractor designed to convert PDF files into structured data,… funstory-ai/babeldoc — BabelDOC is a technical document translation system designed to translate PDF files while preserving their original… facebookresearch/nougat — Nougat is a neural OCR system and LLM document parser designed to convert images of academic PDF documents into… microsoft/table-transformer — Table Transformer is a deep learning framework designed for document layout analysis and the automated extraction of…