awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

20 مستودعات

Awesome GitHub RepositoriesColumnar Formats

Standardized memory layouts for columnar data representation.

Distinct from In-Memory Data Stores: Distinct from general in-memory stores: focuses on the specific columnar memory format standard rather than general volatile storage.

Explore 20 awesome GitHub repositories matching data & databases · Columnar Formats. Refine with filters or upvote what's useful.

Awesome Columnar Formats GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • bokeh/bokehالصورة الرمزية لـ bokeh

    bokeh/bokeh

    20,403عرض على GitHub↗

    Bokeh is a Python data visualization library and interactive plotting framework used to create high-performance graphics and data dashboards that render in web browsers. It serves as a tool for generating standalone HTML documents, embedded components for digital notebooks, and full-stack web applications powered by a Python backend. The project distinguishes itself through its ability to handle large or streaming datasets while maintaining smooth interactivity. It enables linked brushing across multiple views, allowing data selected in one plot to automatically highlight corresponding data i

    Organizes datasets into named arrays using a columnar format to optimize data transfer from Python to the browser.

    TypeScriptbokehdata-visualisationinteractive-plots
    عرض على GitHub↗20,403
  • apache/arrowالصورة الرمزية لـ apache

    apache/arrow

    16,529عرض على GitHub↗

    Arrow is a cross-language development platform for in-memory data. It provides a standardized, language-independent columnar memory format designed to accelerate analytical operations and improve memory efficiency on modern computing hardware. By utilizing a schema-driven approach, the framework enables the efficient organization of both flat and nested data structures. The project functions as an analytical data processing engine that facilitates high-performance computation directly on memory-resident datasets. It distinguishes itself through a zero-copy architecture, which allows multiple

    Provides a language-independent standard for organizing flat or nested data in memory.

    C++arrowparquet
    عرض على GitHub↗16,529
  • jpmorganchase/python-trainingالصورة الرمزية لـ jpmorganchase

    jpmorganchase/python-training

    12,714عرض على GitHub↗

    This project is a comprehensive educational curriculum designed to teach Python programming through the lens of data science and financial analysis. It provides a structured guide for learning how to process complex numerical information, build data models, and perform scientific computing tasks using standard industry libraries. The materials focus on practical applications, enabling users to develop skills in financial data analysis and interactive exploration. By working through these resources, learners gain experience in executing high-performance mathematical operations, transforming ra

    Organizes tabular information into contiguous memory blocks for rapid filtering and aggregation.

    Jupyter Notebookbankingbinderbinder-ready
    عرض على GitHub↗12,714
  • popcorn-official/popcorn-desktopالصورة الرمزية لـ popcorn-official

    popcorn-official/popcorn-desktop

    10,605عرض على GitHub↗

    Popcorn Desktop is a multi-platform video aggregator and BitTorrent media streamer that provides a unified interface for watching movies and television series. It functions as a local media player capable of rendering video files stored directly on the host file system. The application distinguishes itself by acting as a content dataset publisher, exporting large movie and series catalogs in columnar formats to facilitate analysis by external researchers. Its broader capabilities include aggregated content federation and stream routing across multiple platforms, allowing for the integration

    Transforms internal content catalogs into column-oriented formats to facilitate high-performance analysis by external researchers.

    TypeScript
    عرض على GitHub↗10,605
  • dr5hn/countries-states-cities-databaseالصورة الرمزية لـ dr5hn

    dr5hn/countries-states-cities-database

    9,291عرض على GitHub↗

    This project is a comprehensive geographic location dataset and reference library providing standardized data for countries, states, and cities. It serves as a source of truth for regional hierarchies, ISO codes, coordinates, and timezone information, available as both a relational SQL database and a document-based JSON library. The project includes a custom dataset export tool that functions as a filtering engine. This allows for the generation of tailored geographic files in JSON, CSV, and GeoJSON formats by selecting only the specific regions or fields required. The dataset covers global

    Provides a browser-based tool to export tailored datasets by selecting specific countries, regions, and fields.

    Pythoncitiescountriescountry
    عرض على GitHub↗9,291
  • a16z-infra/ai-townالصورة الرمزية لـ a16z-infra

    a16z-infra/ai-town

    9,285عرض على GitHub↗

    AI Town is a TypeScript-based simulation engine used to create virtual environments where autonomous characters interact and socialize. It functions as a framework for orchestrating multiple AI agents within a persistent digital world, utilizing language models and a game engine to drive character behavior and social interactions. The project differentiates itself through a dedicated agent sandbox and a vector database agent store, which allow for the management of agent memories and world state. It integrates generative AI for background music and provides tools for simulation world design,

    Enables exporting table data and file storage content to local files for external use.

    TypeScript
    عرض على GitHub↗9,285
  • apify/crawlee-pythonالصورة الرمزية لـ apify

    apify/crawlee-python

    8,097عرض على GitHub↗

    Crawlee-python is a web crawling framework for building scalable scrapers using Python. It serves as a comprehensive tool for web scraping automation, providing a system to extract structured data from websites using both lightweight HTTP requests and headless browser automation. The framework is distinguished by its anti-bot evasion capabilities, which include browser fingerprint impersonation and tiered proxy rotation to bypass detection systems and solve challenges such as Cloudflare. It also incorporates artificial intelligence for autonomous website navigation and schema-based data extra

    Saves the entire contents of a dataset into a single file in CSV or JSON format.

    Pythonapifyautomationbeautifulsoup
    عرض على GitHub↗8,097
  • jerbouma/financedatabaseالصورة الرمزية لـ JerBouma

    JerBouma/FinanceDatabase

    6,987عرض على GitHub↗

    FinanceDatabase is a system of data repositories and interfaces providing a corporate fundamental database, a financial market data API, and an SEC filings aggregator. It functions as a financial valuation engine and a macroeconomic indicator feed, offering a programmatic way to access market quotes, corporate fundamentals, and official regulatory disclosures. The project distinguishes itself through an institutional ownership tracker that monitors fund holdings, insider trading activity, and political financial disclosures. It also includes a dedicated tool for extracting and analyzing offic

    Retrieves large datasets for company profiles and financial statements in single batch operations.

    Pythonanalysiscryptocurrenciescurrencies
    عرض على GitHub↗6,987
  • lance-format/lanceL

    lance-format/lance

    6,699عرض على GitHub↗

    Lance is a columnar data format and storage layer designed for high-performance random access and the persistence of multimodal data. It functions as a vector database storage system, a multimodal data store, and a versioned dataset manager. The project distinguishes itself as a hybrid search engine that combines vector similarity search and full-text indexing on a single dataset. It provides unified storage for diverse data types including images, audio, and video, utilizing a system that lazy-loads large binary objects only when requested. The system manages dataset evolution through schem

    Provides a storage standard based on Apache Arrow for high-performance random access.

    Rust
    عرض على GitHub↗6,699
  • apache/pinotالصورة الرمزية لـ apache

    apache/pinot

    6,098عرض على GitHub↗

    Pinot is a distributed, columnar analytical database designed for high-concurrency, low-latency query processing. It functions as a real-time OLAP datastore, enabling interactive, user-facing analytics by ingesting and querying massive datasets from both streaming and batch sources. The system architecture relies on a centralized controller for cluster coordination and a distributed segment-based storage model to ensure horizontal scalability. The platform distinguishes itself through a hybrid ingestion pipeline that unifies real-time event streams and historical batch data into a single quer

    Organizes table data into independent, time-based horizontal shards using columnar storage and indices.

    Java
    عرض على GitHub↗6,098
  • awslabs/gluontsالصورة الرمزية لـ awslabs

    awslabs/gluonts

    5,199عرض على GitHub↗

    GluonTS هي مكتبة سلاسل زمنية احتمالية وإطار عمل للتنبؤ بالتعلم العميق. توفر مجموعة أدوات لبناء وتدريب وتقييم بنى الشبكات العصبية التي تتنبأ بالقيم المستقبلية كتوزيعات احتمالية لتحديد عدم اليقين. يتميز المشروع بدعم التنبؤ بدون تدريب مسبق (zero-shot) ودمج نهج نمذجة متنوعة، بما في ذلك الشبكات العصبية الاحتمالية العميقة وأغلفة للمكتبات الإحصائية الخارجية مثل Prophet و R forecast. ينفذ بدائيات معمارية متخصصة مثل الالتفافات السببية والشبكات المتبقية القابلة للعكس لمنع تسرب المعلومات وتعيين التمثيلات الكامنة في توزيعات احتمالية صالحة. يغطي إطار العمل سطح هندسة بيانات شاملاً، بما في ذلك توسيع السلاسل الزمنية، والتحويلات التقابلية، والنمذجة الهرمية. يستخدم Apache Arrow و Parquet لبث مجموعة البيانات عالي الأداء وإدارة الوصول العشوائي. لتقييم النموذج، يتضمن جناح تقييم لقياس دقة التنبؤ والتغطية الاحتمالية باستخدام مقاييس مثل خسارة الكمية ودرجات رتبة الاحتمال المستمرة. تدعم المكتبة نشر النموذج من خلال التكامل مع Amazon SageMaker.

    Converts encoded Apache Arrow data batches into time series formats using predefined schemas.

    Pythonartificial-intelligenceawsdata-science
    عرض على GitHub↗5,199
  • awslabs/gluon-tsالصورة الرمزية لـ awslabs

    awslabs/gluon-ts

    5,200عرض على GitHub↗

    GluonTS هو إطار عمل للتنبؤ بالسلاسل الزمنية الاحتمالية، مصمم للتنبؤ بالقيم المستقبلية كتوزيعات احتمالية مع فترات ثقة. يدعم كلاً من تدريب النموذج التقليدي والتنبؤ بدون تدريب مسبق (zero-shot)، حيث تولد النماذج المدربة مسبقاً تنبؤات لسلاسل جديدة دون تدريب إضافي. يتميز المشروع بدمج مجموعة واسعة من نهج التنبؤ في سير عمل موحد. يتضمن ذلك بنى التعلم العميق مثل الشبكات العصبية المتكررة والالتفافات السببية، بالإضافة إلى دمج النماذج الإحصائية الخارجية، ومكتبة Prophet، وحزم R. توفر مجموعة الأدوات سطحاً شاملاً لهندسة بيانات السلاسل الزمنية، وتغطي توسيع مجموعة البيانات، والتقسيم، وتحويل البيانات الزمنية الخام إلى موترات (tensors). كما تتضمن مجموعة من أدوات التقييم لقياس دقة التنبؤ وفترات عدم اليقين، بالإضافة إلى أدوات لاستمرارية مجموعة البيانات باستخدام تنسيقات مثل Arrow و Parquet. يدعم إطار العمل نشر نماذج التنبؤ داخل البنية التحتية السحابية.

    Implements utilities for exporting structured datasets into high-performance columnar formats like Feather and Parquet.

    Python
    عرض على GitHub↗5,200
  • rom1504/img2datasetالصورة الرمزية لـ rom1504

    rom1504/img2dataset

    4,423عرض على GitHub↗

    img2dataset هو خط معالجة عالي الأداء لمجموعات بيانات الصور وأداة معالجة مسبقة مصممة لتنزيل ومعالجة ملايين الصور من روابط URL لتدريب تعلم الآلة. يعمل كأداة تنزيل صور موزعة ومصدر بيانات للتخزين السحابي، حيث ينقل مجموعات البيانات البصرية الكبيرة من مصادر الويب مباشرة إلى تنسيقات مهيكلة. يعطي النظام الأولوية للحصول على البيانات ذات الإنتاجية العالية من خلال توزيع أحمال العمل عبر أنوية CPU متعددة وأجهزة متعددة. يتكامل مباشرة مع حاويات التخزين السحابي البعيدة ويستخدم نظام تتبع قائماً على البيان (manifest-based) لاستئناف التنزيلات المتقطعة دون إعادة معالجة البيانات الموجودة. توفر الأداة مجموعة معالجة مسبقة كاملة لإعداد مجموعات بيانات تعلم الآلة، بما في ذلك تغيير حجم الصور، والقص، وترشيح الخصائص بناءً على الحجم أو نسبة العرض إلى الارتفاع. كما تتحقق من سلامة الصور عبر مقارنة الهاش (hash) وتضمن الامتثال لتوجيهات الروبوتات أثناء سير عمل الكشط (scraping). تم تنفيذ المشروع بلغة Python.

    Exports processed images and metadata into specialized file formats optimized for machine learning tools.

    Pythonbig-datadatasetdeep-learning
    عرض على GitHub↗4,423
  • alipay/furyالصورة الرمزية لـ alipay

    alipay/fury

    4,412عرض على GitHub↗

    Fury هو إطار عمل تسلسلي ثنائي متعدد اللغات مصمم لتشفير كائنات المجال والرسوم البيانية المعقدة لتسهيل تبادل البيانات عبر اللغات. يتضمن مترجم لغة تعريف الواجهة (IDL) الذي يترجم تعريفات المخطط إلى أنواع أصلية اصطلاحية ونصوص تسلسلية عبر لغات متعددة. يتميز المشروع بقارئ ثنائي بدون نسخ (zero-copy) يسمح بالوصول إلى حقول محددة دون إلغاء تسلسل الكائن بالكامل، بالإضافة إلى مسلسل رسوم بيانية للكائنات يحافظ على المراجع الدائرية وسلامة المراجع. كما يتميز بمحول بيانات يحول البيانات الثنائية القائمة على الصفوف إلى تنسيقات Apache Arrow القائمة على الأعمدة لأحمال العمل التحليلية. يغطي إطار العمل مجالات قدرة واسعة بما في ذلك تطور المخطط القائم على البيانات الوصفية للتوافق للأمام وللخلف، وعملية تجميع AOT في وقت البناء للقضاء على الانعكاس في وقت التشغيل، وإلغاء التسلسل الآمن عبر التحقق من النوع القائم على القائمة البيضاء. كما يوفر تكاملاً لاستدعاءات الإجراءات عن بُعد عالية الأداء من خلال gRPC.

    Transforms row-based binary data into Apache Arrow RecordBatch formats for high-performance analytical workloads.

    Java
    عرض على GitHub↗4,412
  • apache/lucene-solrالصورة الرمزية لـ apache

    apache/lucene-solr

    4,357عرض على GitHub↗

    هذا المشروع عبارة عن محرك بحث نصي كامل وبنية تحتية للبحث المؤسسي مصممة لفهرسة واسترجاع مجموعات كبيرة من المستندات. يوفر إطار عمل شاملاً لاكتشاف المعلومات باستخدام نتائج مرتبة وتحليل لغوي. يدمج النظام بحث التشابه المتجهي عالي الأبعاد للاسترجاع الدلالي إلى جانب قدرات البحث النصي الكامل التقليدية. يتميز بدعم استرجاع البيانات الجغرافية المكانية، ومعالجة النصوص متعددة اللغات، وسير عمل اقتراحات البحث الذي يتضمن إكمال الاستعلام المتسامح مع الأخطاء الإملائية والتدقيق الإملائي. تغطي المنصة مجموعة واسعة من قدرات البحث والفهرسة، بما في ذلك تنفيذ الاستعلامات المعقدة، وتجميع عدد الأوجه، وتجميع النتائج. يتعامل النظام مع تحليل النصوص من خلال التجزئة والتطبيع، مع تقديم أدوات متخصصة لربط المستندات، وإبراز نتائج البحث، والتسجيل المخصص بناءً على الحداثة والمسافة. تتوفر واجهة بحث بلغة Python لكشف وظائف الفهرسة والاستعلام لبيئات البرمجة الخارجية.

    Provides field data in a column-oriented format to enable fast sorting, faceting, and aggregation.

    backendinformation-retrievaljava
    عرض على GitHub↗4,357
  • facebookincubator/veloxالصورة الرمزية لـ facebookincubator

    facebookincubator/velox

    4,155عرض على GitHub↗

    Velox هو محرك تنفيذ استعلامات عالي الأداء ومكتبة لمعالجة البيانات العمودية بلغة C++. يعمل كإطار عمل قابل للتركيب لتنفيذ محركات الاستعلام التحليلية، ويوفر مقيماً للتعبيرات المتجهة (vectorized) ومجموعة أدوات لأنظمة إدارة البيانات. يتميز المشروع باستخدامه للتنفيذ العمودي المتجه وتخصيص الذاكرة القائم على الساحة (arena-based) لمعالجة مجموعات البيانات واسعة النطاق. يتميز بتحسينات متخصصة مثل التخزين المؤقت لجدول الربط الإذاعي (broadcast join)، ودفع الفلتر الديناميكي للأسفل، وترميز القاموس لتقليل حمل الذاكرة وتسريع القراءات التحليلية. يغطي المحرك مجموعة واسعة من القدرات التحليلية، بما في ذلك تنفيذ عمليات الربط (hash, merge, semi joins)، بالإضافة إلى التجميع المتوازي متعدد المراحل وحساب دوال النافذة. يوفر بدائيات للتخزين العمودي في الذاكرة، وفك تشفير بيانات Parquet، والتكامل مع التخزين السحابي. يتم توفير القابلية للتوسع من خلال نظام تسجيل الدوال للدوال العددية والتجميعية المخصصة، مع توفر روابط عالية المستوى لربط منطق C++ بلغة Python.

    Organizes data using a columnar in-memory layout that supports both scalar and complex nested types.

    C++
    عرض على GitHub↗4,155
  • mbloch/mapshaperالصورة الرمزية لـ mbloch

    mbloch/mapshaper

    4,133عرض على GitHub↗

    Mapshaper هي أداة لمعالجة وتبسيط وتحويل البيانات المتجهة الجغرافية، متاحة كواجهة سطر أوامر، وأداة متصفح ويب، ومكتبة Node.js. تعمل كمسقط للإحداثيات، ومحول للبيانات المتجهة، ومحسن لأصول خرائط الويب مصمم لتحويل مجموعات البيانات المكانية بين أنظمة مرجعية إحداثية وتنسيقات ملفات مختلفة. يتميز المشروع بتبسيط الهندسة مع الحفاظ على الطوبولوجيا، مما يقلل من عدد الرؤوس مع الحفاظ على الحدود المشتركة لمنع الفجوات والتداخلات. كما يعمل على تحسين الأصول للويب من خلال تكميم الإحداثيات وتصفية السمات لتقليل أحجام الملفات. يغطي النظام مجموعة واسعة من الإمكانيات، بما في ذلك إعادة إسقاط الإحداثيات باستخدام سلاسل PROJ ورموز EPSG، وتحويل البيانات عبر تنسيقات مثل Shapefile وGeoJSON وTopoJSON وGeoPackage وKML. ويوفر أدوات معالجة هندسية واسعة النطاق للتخزين المؤقت، والقص، والإذابة، وإصلاح الطوبولوجيا، بالإضافة إلى أدوات إدارة البيانات لربط السمات وتصفيتها وتحويلها. بالإضافة إلى ذلك، يتضمن ميزات تصور لتوليد صادرات SVG مصممة، وشبكات إحداثيات، وخرائط رموز متناسبة. يمكن دمج إمكانيات المعالجة المكانية مباشرة في تطبيقات JavaScript وخطوط أنابيب البناء عبر مكتبة Node.js الخاصة به.

    Reads and writes GeoParquet columnar files to exchange vector geometries and tabular attributes.

    JavaScript
    عرض على GitHub↗4,133
  • uptrace/uptraceالصورة الرمزية لـ uptrace

    uptrace/uptrace

    4,098عرض على GitHub↗

    Uptrace is an OpenTelemetry-based observability platform designed to collect, store, and analyze distributed traces, metrics, and logs. It functions as a centralized logging backend, a distributed tracing system, and a metrics engine to monitor application performance and system health. The platform is distinguished by AI-powered operational capabilities, allowing users to query telemetry data and manage monitoring dashboards using natural language. It specifically includes specialized monitoring for generative AI pipelines, tracking token usage and response quality for LLM interactions and r

    Utilizes OTel Arrow to transport telemetry data in a columnar layout for reduced bandwidth and faster ingestion.

    Goapmapplication-monitoringclickhouse
    عرض على GitHub↗4,098
  • chartbrew/chartbrewالصورة الرمزية لـ chartbrew

    chartbrew/chartbrew

    3,641عرض على GitHub↗

    Chartbrew is a self-hosted business intelligence platform and data visualization engine designed to transform raw data from SQL databases and external API endpoints into interactive charts and dashboards. It serves as a tool for building analytics dashboards that monitor business metrics and KPIs through a privately hosted environment. The platform distinguishes itself with an embedded analytics workflow, allowing users to generate secure, time-limited shared links and iframes to display private charts on external websites. It also provides programmatic chart generation via API and integrates

    Streamlines setup by generating a reusable dataset and its associated data requests in a single operation.

    JavaScriptanalyticsapichartjs
    عرض على GitHub↗3,641
  • google/osv.devالصورة الرمزية لـ google

    google/osv.dev

    2,494عرض على GitHub↗

    OSV is a distributed database and aggregator of open-source security advisories that uses a standardized vulnerability schema to track security flaws. It functions as a system for collecting and normalizing security data from diverse ecosystems into a single unified format, providing a web API for querying package vulnerabilities and submitting standardized records. The project distinguishes itself through a security advisory distribution service that supports bulk dataset exports via cloud storage buckets and incremental synchronization of security record updates. It also employs sandbox-bas

    Exports bulk vulnerability datasets as compressed archives stored in public cloud buckets for high-throughput offline downloading.

    Pythonsecuritysecurity-toolsvulnerability
    عرض على GitHub↗2,494
  1. Home
  2. Data & Databases
  3. In-Memory Data Stores
  4. Columnar Formats

استكشف الوسوم الفرعية

  • Apache Arrow-Based Formats1 وسم فرعيStorage standards specifically leveraging the Apache Arrow memory layout for disk persistence. **Distinct from Columnar Formats:** More specific than general columnar formats, focusing on the Arrow-based interoperability standard.
  • Bulk Dataset ExportUtilities for exporting the entire contents of a structured dataset into a single portable file. **Distinct from Dataset Exporters:** Focuses on bulk archival export of a whole dataset rather than columnar format optimization.
  • Columnar SegmentsIndependent, time-based horizontal shards using columnar storage and indices for memory-mapped query serving. **Distinct from Columnar Formats:** Distinct from general columnar formats: focuses on the segment-based sharding architecture for memory-mapped access.
  • Custom Dataset Generators1 وسم فرعيTools that generate tailored data files from a larger source based on specific constraints. **Distinct from Dataset Exporters:** Distinct from Dataset Exporters by emphasizing the customization and filtering of the output.
  • Dataset ExportersUtilities for exporting structured data into columnar file formats for analysis. **Distinct from Columnar Formats:** Focuses on the export process for researchers rather than the internal memory layout of columnar formats.
  • Row-to-Columnar Conversions1 وسم فرعيProcesses that transform row-oriented data layouts into columnar memory formats. **Distinct from Columnar Formats:** Focuses on the conversion process from row-based binary to columnar format, rather than just the final memory layout.