awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

25 repositorios

Awesome GitHub RepositoriesData Loading Extraction

Libraries for loading, collecting, and extracting data from a variety of data sources and formats.

Explore 25 awesome GitHub repositories matching part of an awesome list · Data Loading Extraction. Refine with filters or upvote what's useful.

Awesome Data Loading Extraction GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • huggingface/datasetsAvatar de huggingface

    huggingface/datasets

    21,643Ver en GitHub↗

    Datasets is a library designed for the management, processing, and sharing of large-scale data collections for machine learning workflows. It functions as both a data processing framework and a versioning platform, providing tools to organize, filter, and transform massive datasets while ensuring reproducibility across research and development teams. The library distinguishes itself by enabling the handling of datasets that exceed available system memory. It utilizes memory-mapped file access, disk-based caching, and lazy iterative streaming to maintain performance when working with large-sca

    Hub of ready-to-use datasets for AI.

    Pythonaiartificial-intelligencecomputer-vision
    Ver en GitHub↗21,643
  • joke2k/fakerAvatar de joke2k

    joke2k/faker

    19,278Ver en GitHub↗

    Faker is a Python library designed to generate realistic synthetic data for software testing, database prototyping, and privacy-preserving anonymization. It provides a comprehensive suite of tools to create diverse information types, including personal identities, financial records, geographic locations, and technical system metadata, allowing developers to populate environments with mock data that mimics real-world structures. The library is built on a modular provider architecture that supports dynamic method dispatch, enabling users to extend functionality by registering custom data genera

    Generates fake data.

    Pythondatasetfakefake-data
    Ver en GitHub↗19,278
  • wireservice/csvkitAvatar de wireservice

    wireservice/csvkit

    6,390Ver en GitHub↗

    csvkit is a composable Unix-style command-line toolkit for converting, filtering, and analyzing CSV files directly from the terminal. It provides a suite of focused single-purpose commands that can be combined via pipes to build complex data processing workflows, with a modular architecture that includes a column-type inference engine for automatically detecting data types and a streaming-pipeline design for efficient handling of tabular data. The toolkit distinguishes itself through its SQL-engine abstraction layer, which allows users to run SQL queries directly against CSV files without req

    Utilities for converting and working with CSV.

    Python
    Ver en GitHub↗6,390
  • snorkel-team/snorkelAvatar de snorkel-team

    snorkel-team/snorkel

    5,981Ver en GitHub↗

    Snorkel is a weak supervision system that enables users to programmatically generate training labels for machine learning models without manual annotation. At its core, it provides a framework for writing labeling functions as Python callables that each vote on data points, and then trains a probabilistic graphical model over these multiple weak supervision sources to estimate latent true labels without any ground truth data. The system automatically learns accuracy and correlation parameters between labeling functions by analyzing observed agreement patterns on unlabeled data, converting lab

    Generating training data with weak supervision.

    Python
    Ver en GitHub↗5,981
  • martinblech/xmltodictAvatar de martinblech

    martinblech/xmltodict

    5,741Ver en GitHub↗

    xmltodict es una librería de Python que proporciona serialización bidireccional entre documentos XML y diccionarios. Funciona como un analizador (parser) que convierte entradas marcadas en pares clave-valor y como una utilidad de serialización que transforma diccionarios de nuevo en documentos XML estructurados. El proyecto incluye un procesador de flujo incremental que utiliza callbacks basados en profundidad para manejar archivos XML grandes manteniendo un uso de memoria constante. Cuenta con un gestor de espacios de nombres (namespaces) para asignar prefijos y declaraciones, así como un sanitizador de seguridad que bloquea la expansión de entidades externas y valida nombres de elementos para prevenir ataques de inyección. La librería proporciona capacidades para la aplicación de tipos de datos, como forzar que elementos específicos se representen como listas independientemente del número de hijos. También admite el post-procesamiento de datos a través de callbacks definidos por el usuario y ofrece controles configurables para expandir, colapsar u omitir espacios de nombres durante el proceso de conversión.

    Makes XML feel like working with JSON.

    Python
    Ver en GitHub↗5,741
  • wkentaro/gdownAvatar de wkentaro

    wkentaro/gdown

    5,116Ver en GitHub↗

    gdown is a command-line tool that downloads public files and folders from Google Drive without requiring authentication. It bypasses the mandatory virus-scan warning page to retrieve large files that conventional download tools block, and can resume interrupted transfers using HTTP range requests. Beyond simple file downloads, gdown can recursively download entire folder hierarchies while preserving the local directory structure. It lists the contents of a public folder as structured JSON without downloading the files themselves, and resolves a file's real name and extension without retrievin

    Google Drive public file downloader.

    Pythoncurldownloaddownloader
    Ver en GitHub↗5,116
  • jazzband/tablibAvatar de jazzband

    jazzband/tablib

    4,754Ver en GitHub↗

    Tablib es una biblioteca de Python diseñada para importar, exportar y manipular conjuntos de datos tabulares. Funciona como un convertidor y gestor de datos multiformato, permitiendo a los usuarios mover información entre diferentes estándares de archivos. La biblioteca admite la transformación de datos a través de formatos CSV, JSON, YAML y Excel. Proporciona una interfaz programática para gestionar estos conjuntos de datos añadiendo filas, filtrando columnas y segregando registros. El sistema utiliza una representación interna común y un mapeo basado en adaptadores para normalizar diversas fuentes de entrada. Esto permite rutinas de lectura y escritura consistentes en todos los formatos de archivo admitidos.

    Tabular dataset module for multiple formats.

    Python
    Ver en GitHub↗4,754
  • deanmalmgren/textractAvatar de deanmalmgren

    deanmalmgren/textract

    4,623Ver en GitHub↗

    Textract is a multi-format text extraction tool and parser. It provides a unified interface to extract plain text from a variety of sources, including documents, images, and audio files. The system functions as a document content parser for PDFs and spreadsheets, an image text extractor using optical character recognition, and a speech-to-text transcriber for audio recordings.

    Extract text from any document.

    HTML
    Ver en GitHub↗4,623
  • rom1504/img2datasetAvatar de rom1504

    rom1504/img2dataset

    4,423Ver en GitHub↗

    img2dataset es un pipeline de datasets de imágenes de alto rendimiento y herramienta de preprocesamiento diseñada para descargar y procesar millones de imágenes desde URLs para el entrenamiento de machine learning. Funciona como un descargador de imágenes distribuido y exportador de datos a almacenamiento en la nube, moviendo grandes datasets visuales desde fuentes web directamente a formatos estructurados. El sistema prioriza la adquisición de datos de alto rendimiento distribuyendo cargas de trabajo entre múltiples núcleos de CPU y máquinas. Se integra directamente con buckets de almacenamiento en la nube remotos y emplea un sistema de seguimiento basado en manifiestos para reanudar descargas interrumpidas sin reprocesar los datos existentes. La herramienta proporciona una suite de preprocesamiento completa para la preparación de datasets de machine learning, incluyendo redimensionamiento de imágenes, recorte y filtrado de propiedades basado en tamaño o relación de aspecto. También verifica la integridad de la imagen mediante comparación de hash y asegura el cumplimiento de las directivas de robots durante el flujo de trabajo de scraping. El proyecto está implementado en Python.

    Turn image URLs into datasets.

    Pythonbig-datadatasetdeep-learning
    Ver en GitHub↗4,423
  • camelot-dev/camelotAvatar de camelot-dev

    camelot-dev/camelot

    3,764Ver en GitHub↗

    Camelot is a Python library and processing engine designed to extract tabular data from PDF documents. It converts unstructured tables into machine-readable formats such as CSV, JSON, and Excel. The project provides specialized toolsets for different document types, using line detection for ruled tables and whitespace analysis for borderless tables. It includes an optical character recognition system to recover structured data from image-based scanned PDFs that lack a digital text layer. The library handles complex document layouts, including encrypted files, rotated pages, and tables that s

    Extract tabular data from PDFs.

    Python
    Ver en GitHub↗3,764
  • sdv-dev/sdvAvatar de sdv-dev

    sdv-dev/SDV

    3,508Ver en GitHub↗

    Synthetic data generation for tabular data

    Synthetic data generation for tabular data.

    Python
    Ver en GitHub↗3,508
  • xlwings/xlwingsAvatar de xlwings

    xlwings/xlwings

    3,312Ver en GitHub↗

    xlwings - Make Excel fly with Python!

    Call Python from Excel.

    Pythonautomationexcelgoogle-sheets
    Ver en GitHub↗3,312
  • pydata/pandas-datareaderAvatar de pydata

    pydata/pandas-datareader

    3,217Ver en GitHub↗

    Extract data from a wide range of Internet sources into a pandas DataFrame.

    Extract data from internet sources.

    Pythondatadata-analysisdataset
    Ver en GitHub↗3,217
  • ahupp/python-magicAvatar de ahupp

    ahupp/python-magic

    2,886Ver en GitHub↗

    python-magic is a C-binding wrapper that provides a Python interface for the libmagic system library. It functions as a file signature analyzer and MIME type detector, identifying file formats by comparing header bytes against a database of known binary signatures. The library enables the identification of file types from both file paths and raw data buffers. It supports custom file signature matching through the injection of user-provided magic databases, allowing for the detection of specialized or proprietary formats. The project covers binary data analysis and MIME type mapping to transl

    Wrapper for libmagic.

    Python
    Ver en GitHub↗2,886
  • python-excel/xlrdAvatar de python-excel

    python-excel/xlrd

    2,205Ver en GitHub↗

    Please use openpyxl where you can...

    Reading Excel files.

    Python
    Ver en GitHub↗2,205
  • camelot-dev/excaliburAvatar de camelot-dev

    camelot-dev/excalibur

    1,790Ver en GitHub↗

    Web interface for PDF table extraction.

    Pythonextractfor-humanspdf
    Ver en GitHub↗1,790
  • singer-io/getting-startedAvatar de singer-io

    singer-io/getting-started

    1,342Ver en GitHub↗

    Singer is an open source standard for moving data between databases, web APIs, files, queues, and just about anything else you can think of. The Singer spec describes how data extraction scripts — called “Taps” — and data loading scripts — called “Targets” — should communicate using a standard…

    Standard for moving data between systems.

    Makefile
    Ver en GitHub↗1,342
  • intake/intakeAvatar de intake

    intake/intake

    1,080Ver en GitHub↗

    Intake is a lightweight package for finding, investigating, loading and disseminating data.

    Package for finding and loading data.

    Python
    Ver en GitHub↗1,080
  • simonw/csvs-to-sqliteAvatar de simonw

    simonw/csvs-to-sqlite

    932Ver en GitHub↗

    Convert CSV files into a SQLite database. Browse and publish that SQLite database with Datasette.

    Convert CSV to SQLite.

    Python
    Ver en GitHub↗932
  • upgini/upginiAvatar de upgini

    upgini/upgini

    351Ver en GitHub↗

    Data search & enrichment library for Machine Learning → Easily find and add relevant features to your ML & AI pipeline from hundreds of public and premium external data sources, including open & commercial LLMs

    Data search and enrichment for ML.

    Python
    Ver en GitHub↗351
Ant.12Siguiente
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Data Loading Extraction