awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

20 repository-uri

Awesome GitHub RepositoriesColumnar Formats

Standardized memory layouts for columnar data representation.

Distinct from In-Memory Data Stores: Distinct from general in-memory stores: focuses on the specific columnar memory format standard rather than general volatile storage.

Explore 20 awesome GitHub repositories matching data & databases · Columnar Formats. Refine with filters or upvote what's useful.

Awesome Columnar Formats GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • bokeh/bokehAvatar bokeh

    bokeh/bokeh

    20,403Vezi pe GitHub↗

    Bokeh is a Python data visualization library and interactive plotting framework used to create high-performance graphics and data dashboards that render in web browsers. It serves as a tool for generating standalone HTML documents, embedded components for digital notebooks, and full-stack web applications powered by a Python backend. The project distinguishes itself through its ability to handle large or streaming datasets while maintaining smooth interactivity. It enables linked brushing across multiple views, allowing data selected in one plot to automatically highlight corresponding data i

    Organizes datasets into named arrays using a columnar format to optimize data transfer from Python to the browser.

    TypeScriptbokehdata-visualisationinteractive-plots
    Vezi pe GitHub↗20,403
  • apache/arrowAvatar apache

    apache/arrow

    16,529Vezi pe GitHub↗

    Arrow is a cross-language development platform for in-memory data. It provides a standardized, language-independent columnar memory format designed to accelerate analytical operations and improve memory efficiency on modern computing hardware. By utilizing a schema-driven approach, the framework enables the efficient organization of both flat and nested data structures. The project functions as an analytical data processing engine that facilitates high-performance computation directly on memory-resident datasets. It distinguishes itself through a zero-copy architecture, which allows multiple

    Provides a language-independent standard for organizing flat or nested data in memory.

    C++arrowparquet
    Vezi pe GitHub↗16,529
  • jpmorganchase/python-trainingAvatar jpmorganchase

    jpmorganchase/python-training

    12,714Vezi pe GitHub↗

    This project is a comprehensive educational curriculum designed to teach Python programming through the lens of data science and financial analysis. It provides a structured guide for learning how to process complex numerical information, build data models, and perform scientific computing tasks using standard industry libraries. The materials focus on practical applications, enabling users to develop skills in financial data analysis and interactive exploration. By working through these resources, learners gain experience in executing high-performance mathematical operations, transforming ra

    Organizes tabular information into contiguous memory blocks for rapid filtering and aggregation.

    Jupyter Notebookbankingbinderbinder-ready
    Vezi pe GitHub↗12,714
  • popcorn-official/popcorn-desktopAvatar popcorn-official

    popcorn-official/popcorn-desktop

    10,605Vezi pe GitHub↗

    Popcorn Desktop is a multi-platform video aggregator and BitTorrent media streamer that provides a unified interface for watching movies and television series. It functions as a local media player capable of rendering video files stored directly on the host file system. The application distinguishes itself by acting as a content dataset publisher, exporting large movie and series catalogs in columnar formats to facilitate analysis by external researchers. Its broader capabilities include aggregated content federation and stream routing across multiple platforms, allowing for the integration

    Transforms internal content catalogs into column-oriented formats to facilitate high-performance analysis by external researchers.

    TypeScript
    Vezi pe GitHub↗10,605
  • dr5hn/countries-states-cities-databaseAvatar dr5hn

    dr5hn/countries-states-cities-database

    9,291Vezi pe GitHub↗

    This project is a comprehensive geographic location dataset and reference library providing standardized data for countries, states, and cities. It serves as a source of truth for regional hierarchies, ISO codes, coordinates, and timezone information, available as both a relational SQL database and a document-based JSON library. The project includes a custom dataset export tool that functions as a filtering engine. This allows for the generation of tailored geographic files in JSON, CSV, and GeoJSON formats by selecting only the specific regions or fields required. The dataset covers global

    Provides a browser-based tool to export tailored datasets by selecting specific countries, regions, and fields.

    Pythoncitiescountriescountry
    Vezi pe GitHub↗9,291
  • a16z-infra/ai-townAvatar a16z-infra

    a16z-infra/ai-town

    9,285Vezi pe GitHub↗

    AI Town is a TypeScript-based simulation engine used to create virtual environments where autonomous characters interact and socialize. It functions as a framework for orchestrating multiple AI agents within a persistent digital world, utilizing language models and a game engine to drive character behavior and social interactions. The project differentiates itself through a dedicated agent sandbox and a vector database agent store, which allow for the management of agent memories and world state. It integrates generative AI for background music and provides tools for simulation world design,

    Enables exporting table data and file storage content to local files for external use.

    TypeScript
    Vezi pe GitHub↗9,285
  • apify/crawlee-pythonAvatar apify

    apify/crawlee-python

    8,097Vezi pe GitHub↗

    Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive tool for web scraping automation, providing a system to extract structured data from websites using both lightweight HTTP requests and headless browser automation. The framework is distinguished by its anti-bot evasion capabilities, which include browser fingerprint impersonation and tiered proxy rotation to bypass detection systems and solve challenges such as Cloudflare. It also incorporates artificial intelligence for autonomous website navigation and schema-based data extra

    Saves the entire contents of a dataset into a single file in CSV or JSON format.

    Pythonapifyautomationbeautifulsoup
    Vezi pe GitHub↗8,097
  • jerbouma/financedatabaseAvatar JerBouma

    JerBouma/FinanceDatabase

    6,987Vezi pe GitHub↗

    FinanceDatabase is a system of data repositories and interfaces providing a corporate fundamental database, a financial market data API, and an SEC filings aggregator. It functions as a financial valuation engine and a macroeconomic indicator feed, offering a programmatic way to access market quotes, corporate fundamentals, and official regulatory disclosures. The project distinguishes itself through an institutional ownership tracker that monitors fund holdings, insider trading activity, and political financial disclosures. It also includes a dedicated tool for extracting and analyzing offic

    Retrieves large datasets for company profiles and financial statements in single batch operations.

    Pythonanalysiscryptocurrenciescurrencies
    Vezi pe GitHub↗6,987
  • lance-format/lanceL

    lance-format/lance

    6,699Vezi pe GitHub↗

    Lance is a columnar data format and storage layer designed for high-performance random access and the persistence of multimodal data. It functions as a vector database storage system, a multimodal data store, and a versioned dataset manager. The project distinguishes itself as a hybrid search engine that combines vector similarity search and full-text indexing on a single dataset. It provides unified storage for diverse data types including images, audio, and video, utilizing a system that lazy-loads large binary objects only when requested. The system manages dataset evolution through schem

    Provides a storage standard based on Apache Arrow for high-performance random access.

    Rust
    Vezi pe GitHub↗6,699
  • apache/pinotAvatar apache

    apache/pinot

    6,098Vezi pe GitHub↗

    Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It functions as a real-time OLAP datastore, enabling interactive, user-facing analytics by ingesting and querying massive datasets from both streaming and batch sources. The system architecture relies on a centralized controller for cluster coordination and a distributed segment-based storage model to ensure horizontal scalability. The platform distinguishes itself through a hybrid ingestion pipeline that unifies real-time event streams and historical batch data into a single quer

    Organizes table data into independent, time-based horizontal shards using columnar storage and indices.

    Java
    Vezi pe GitHub↗6,098
  • awslabs/gluontsAvatar awslabs

    awslabs/gluonts

    5,199Vezi pe GitHub↗

    GluonTS este o bibliotecă de serii temporale probabilistice și un framework de prognoză prin deep learning. Oferă un toolkit pentru construirea, antrenarea și evaluarea arhitecturilor de rețele neuronale care prezic valori viitoare ca distribuții de probabilitate pentru a cuantifica incertitudinea. Proiectul se distinge prin suportul pentru prognoza zero-shot și integrarea unor abordări de modelare diverse, incluzând rețele neuronale probabilistice profunde și wrapper-e pentru biblioteci statistice externe precum Prophet și R forecast. Implementează primitive arhitecturale specializate precum convoluțiile cauzale și rețelele reziduale inversabile pentru a preveni scurgerea informațiilor și a mapa reprezentările latente în distribuții de probabilitate valide. Framework-ul acoperă o suprafață cuprinzătoare de inginerie a datelor, incluzând scalarea seriilor temporale, transformări bijective și modelare ierarhică. Utilizează Apache Arrow și Parquet pentru streaming-ul seturilor de date de înaltă performanță și gestionarea accesului aleatoriu. Pentru evaluarea modelului, include o suită de evaluare pentru măsurarea acurateței prognozei și a acoperirii probabilistice folosind metrici precum quantile loss și continuous rank probability scores. Biblioteca suportă implementarea modelului prin integrarea cu Amazon SageMaker.

    Converts encoded Apache Arrow data batches into time series formats using predefined schemas.

    Pythonartificial-intelligenceawsdata-science
    Vezi pe GitHub↗5,199
  • awslabs/gluon-tsAvatar awslabs

    awslabs/gluon-ts

    5,200Vezi pe GitHub↗

    GluonTS este un framework pentru prognoza probabilistică a seriilor temporale, conceput pentru a prezice valori viitoare ca distribuții de probabilitate cu intervale de încredere. Suportă atât antrenarea modelelor tradiționale, cât și prognoza zero-shot, unde modelele preantrenate generează predicții pentru serii noi fără antrenare suplimentară. Proiectul se distinge prin integrarea unei mari varietăți de abordări de prognoză într-un flux de lucru unificat. Aceasta include arhitecturi de deep learning precum rețelele neuronale recurente și convoluțiile cauzale, precum și integrarea modelelor statistice externe, a bibliotecii Prophet și a pachetelor R. Toolkit-ul oferă o suprafață cuprinzătoare pentru ingineria datelor de serii temporale, acoperind scalarea seturilor de date, divizarea și transformarea datelor temporale brute în tensori. Include, de asemenea, o suită de instrumente de evaluare pentru măsurarea acurateței prognozei și a intervalelor de incertitudine, precum și utilitare pentru persistența seturilor de date folosind formate precum Arrow și Parquet. Framework-ul suportă implementarea modelelor de prognoză în cadrul infrastructurii cloud.

    Implements utilities for exporting structured datasets into high-performance columnar formats like Feather and Parquet.

    Python
    Vezi pe GitHub↗5,200
  • rom1504/img2datasetAvatar rom1504

    rom1504/img2dataset

    4,423Vezi pe GitHub↗

    img2dataset este un pipeline de seturi de date de imagini de înaltă performanță și un instrument de preprocesare conceput pentru a descărca și procesa milioane de imagini de la URL-uri pentru antrenarea modelelor de machine learning. Funcționează ca un downloader distribuit de imagini și exportator de date în cloud storage, mutând seturi de date vizuale mari din surse web direct în formate structurate. Sistemul prioritizează achiziția de date cu throughput ridicat prin distribuirea sarcinilor pe mai multe nuclee CPU și mașini. Se integrează direct cu bucket-uri de stocare cloud remote și folosește un sistem de urmărire bazat pe manifest pentru a relua descărcările întrerupte fără a reprocesa datele existente. Instrumentul oferă o suită completă de preprocesare pentru pregătirea seturilor de date de machine learning, inclusiv redimensionarea imaginilor, decuparea și filtrarea proprietăților bazată pe dimensiune sau aspect ratio. De asemenea, verifică integritatea imaginilor prin compararea hash-urilor și asigură conformitatea cu directivele roboților în timpul fluxului de lucru de scraping. Proiectul este implementat în Python.

    Exports processed images and metadata into specialized file formats optimized for machine learning tools.

    Pythonbig-datadatasetdeep-learning
    Vezi pe GitHub↗4,423
  • alipay/furyAvatar alipay

    alipay/fury

    4,412Vezi pe GitHub↗

    Fury este un framework de serializare binară multi-limbaj conceput pentru encodarea obiectelor de domeniu și a grafurilor complexe, pentru a facilita schimbul de date între limbaje. Include un compilator de limbaj de definire a interfețelor (IDL) care traduce definițiile de schemă în tipuri native idiomatice și boilerplate de serializare în mai multe limbaje. Proiectul se distinge printr-un cititor binar zero-copy care permite accesarea unor câmpuri specifice fără a deserializa întregul obiect, precum și un serializator de grafuri de obiecte care păstrează referințele circulare și integritatea referențială. De asemenea, dispune de un convertor de date care transformă datele binare bazate pe rânduri în formate coloanare Apache Arrow pentru sarcini analitice. Framework-ul acoperă domenii largi de capabilități, inclusiv evoluția schemei bazată pe metadate pentru compatibilitate înainte și înapoi, un proces de compilare AOT (Ahead-Of-Time) pentru a elimina reflexia la runtime și deserializarea securizată prin validarea tipurilor bazată pe whitelist. Oferă, de asemenea, integrare pentru apeluri de proceduri la distanță de înaltă performanță prin gRPC.

    Transforms row-based binary data into Apache Arrow RecordBatch formats for high-performance analytical workloads.

    Java
    Vezi pe GitHub↗4,412
  • apache/lucene-solrAvatar apache

    apache/lucene-solr

    4,357Vezi pe GitHub↗

    Acest proiect este un motor de căutare full-text și o infrastructură de căutare enterprise concepută pentru indexarea și regăsirea unor seturi mari de documente. Oferă un framework cuprinzător pentru descoperirea informațiilor folosind rezultate ierarhizate și analiză lingvistică. Sistemul integrează căutarea prin similaritate vectorială de înaltă dimensiune pentru regăsirea semantică, alături de capabilitățile tradiționale full-text. Se distinge prin suportul pentru regăsirea datelor geospațiale, procesarea textului multilingv și un flux de lucru pentru sugestii de căutare care include completarea interogărilor cu toleranță la greșeli de scriere și corectarea ortografică. Platforma acoperă o gamă largă de capabilități de căutare și indexare, inclusiv execuția de interogări complexe, agregarea numărului de fațete și gruparea rezultatelor. Gestionează analiza textului prin tokenizare și normalizare, oferind în același timp instrumente specializate pentru unirea documentelor, evidențierea rezultatelor căutării și punctarea personalizată bazată pe recență și distanță. Este disponibilă o interfață de căutare Python pentru a expune funcționalitățile de indexare și interogare către medii de programare externe.

    Provides field data in a column-oriented format to enable fast sorting, faceting, and aggregation.

    backendinformation-retrievaljava
    Vezi pe GitHub↗4,357
  • facebookincubator/veloxAvatar facebookincubator

    facebookincubator/velox

    4,155Vezi pe GitHub↗

    Velox este un motor de execuție a interogărilor C++ de înaltă performanță și o bibliotecă de procesare a datelor coloanare. Servește drept framework compozabil pentru implementarea motoarelor de interogare analitică, oferind un evaluator de expresii vectorizat și un toolkit pentru sistemele de gestionare a datelor. Proiectul se distinge prin utilizarea execuției coloanare vectorizate și a alocării memoriei bazate pe arene pentru a procesa seturi de date la scară largă. Dispune de optimizări specializate, cum ar fi caching-ul tabelelor de broadcast join, push-down dinamic al filtrelor și codificare prin dicționar pentru a reduce overhead-ul de memorie și a accelera citirile analitice. Motorul acoperă o gamă largă de capabilități analitice, inclusiv implementarea de hash, merge și semi joins, precum și agregarea paralelă în mai multe etape și calculul funcțiilor de fereastră. Oferă primitive pentru stocarea coloanară în memorie, decodarea datelor Parquet și integrarea cu stocarea în cloud. Extensibilitatea este oferită printr-un sistem de înregistrare a funcțiilor pentru funcții scalare și agregate personalizate, cu binding-uri de nivel înalt disponibile pentru a conecta logica C++ la Python.

    Organizes data using a columnar in-memory layout that supports both scalar and complex nested types.

    C++
    Vezi pe GitHub↗4,155
  • mbloch/mapshaperAvatar mbloch

    mbloch/mapshaper

    4,133Vezi pe GitHub↗

    Mapshaper este un instrument pentru procesarea, simplificarea și convertirea datelor vectoriale geografice, disponibil ca interfață de linie de comandă, instrument de browser web și bibliotecă Node.js. Funcționează ca un proiector de coordonate, convertor de date vectoriale și optimizator de active pentru hărți web, conceput pentru a transforma seturile de date spațiale între diferite sisteme de referință de coordonate și formate de fișiere. Proiectul se distinge prin simplificarea geometriei care păstrează topologia, ceea ce reduce numărul de noduri (vertex) menținând în același timp limitele partajate pentru a preveni golurile și suprapunerile. Optimizează în continuare activele pentru web prin cuantificarea coordonatelor și filtrarea atributelor pentru a reduce dimensiunile fișierelor. Sistemul acoperă o gamă largă de capabilități, inclusiv reproiectarea coordonatelor folosind șiruri PROJ și coduri EPSG, și conversia datelor între formate precum Shapefile, GeoJSON, TopoJSON, GeoPackage și KML. Oferă instrumente extinse de procesare a geometriei pentru buffering, clipping, dizolvare și repararea topologiilor, precum și utilitare de gestionare a datelor pentru unirea atributelor, filtrare și transformare. În plus, include funcții de vizualizare pentru generarea de exporturi SVG stilizate, graticule și hărți cu simboluri proporționale. Capabilitățile de procesare spațială pot fi integrate direct în aplicațiile JavaScript și în pipeline-urile de build prin biblioteca sa Node.js.

    Reads and writes GeoParquet columnar files to exchange vector geometries and tabular attributes.

    JavaScript
    Vezi pe GitHub↗4,133
  • uptrace/uptraceAvatar uptrace

    uptrace/uptrace

    4,098Vezi pe GitHub↗

    Uptrace is an OpenTelemetry-based observability platform designed to collect, store, and analyze distributed traces, metrics, and logs. It functions as a centralized logging backend, a distributed tracing system, and a metrics engine to monitor application performance and system health. The platform is distinguished by AI-powered operational capabilities, allowing users to query telemetry data and manage monitoring dashboards using natural language. It specifically includes specialized monitoring for generative AI pipelines, tracking token usage and response quality for LLM interactions and r

    Utilizes OTel Arrow to transport telemetry data in a columnar layout for reduced bandwidth and faster ingestion.

    Goapmapplication-monitoringclickhouse
    Vezi pe GitHub↗4,098
  • chartbrew/chartbrewAvatar chartbrew

    chartbrew/chartbrew

    3,641Vezi pe GitHub↗

    Chartbrew is a self-hosted business intelligence platform and data visualization engine designed to transform raw data from SQL databases and external API endpoints into interactive charts and dashboards. It serves as a tool for building analytics dashboards that monitor business metrics and KPIs through a privately hosted environment. The platform distinguishes itself with an embedded analytics workflow, allowing users to generate secure, time-limited shared links and iframes to display private charts on external websites. It also provides programmatic chart generation via API and integrates

    Streamlines setup by generating a reusable dataset and its associated data requests in a single operation.

    JavaScriptanalyticsapichartjs
    Vezi pe GitHub↗3,641
  • google/osv.devAvatar google

    google/osv.dev

    2,494Vezi pe GitHub↗

    OSV is a distributed database and aggregator of open-source security advisories that uses a standardized vulnerability schema to track security flaws. It functions as a system for collecting and normalizing security data from diverse ecosystems into a single unified format, providing a web API for querying package vulnerabilities and submitting standardized records. The project distinguishes itself through a security advisory distribution service that supports bulk dataset exports via cloud storage buckets and incremental synchronization of security record updates. It also employs sandbox-bas

    Exports bulk vulnerability datasets as compressed archives stored in public cloud buckets for high-throughput offline downloading.

    Pythonsecuritysecurity-toolsvulnerability
    Vezi pe GitHub↗2,494
  1. Home
  2. Data & Databases
  3. In-Memory Data Stores
  4. Columnar Formats

Explorează sub-etichetele

  • Apache Arrow-Based Formats1 sub-tagStorage standards specifically leveraging the Apache Arrow memory layout for disk persistence. **Distinct from Columnar Formats:** More specific than general columnar formats, focusing on the Arrow-based interoperability standard.
  • Bulk Dataset ExportUtilities for exporting the entire contents of a structured dataset into a single portable file. **Distinct from Dataset Exporters:** Focuses on bulk archival export of a whole dataset rather than columnar format optimization.
  • Columnar SegmentsIndependent, time-based horizontal shards using columnar storage and indices for memory-mapped query serving. **Distinct from Columnar Formats:** Distinct from general columnar formats: focuses on the segment-based sharding architecture for memory-mapped access.
  • Custom Dataset Generators1 sub-tagTools that generate tailored data files from a larger source based on specific constraints. **Distinct from Dataset Exporters:** Distinct from Dataset Exporters by emphasizing the customization and filtering of the output.
  • Dataset ExportersUtilities for exporting structured data into columnar file formats for analysis. **Distinct from Columnar Formats:** Focuses on the export process for researchers rather than the internal memory layout of columnar formats.
  • Row-to-Columnar Conversions1 sub-tagProcesses that transform row-oriented data layouts into columnar memory formats. **Distinct from Columnar Formats:** Focuses on the conversion process from row-based binary to columnar format, rather than just the final memory layout.