awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 个仓库

Awesome GitHub RepositoriesMemory-Disk Layering

Architectures that layer memory storage over persistent disk storage for performance.

Distinct from Persistent Storage Providers: Focuses on the layering of memory and disk, distinct from general persistence providers.

Explore 14 awesome GitHub repositories matching data & databases · Memory-Disk Layering. Refine with filters or upvote what's useful.

Awesome Memory-Disk Layering GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • pubkey/rxdbpubkey 的头像

    pubkey/rxdb

    23,048在 GitHub 上查看↗

    This project is a reactive, offline-first NoSQL database engine designed for JavaScript applications. It provides a robust framework for managing application state by synchronizing data across browsers, mobile devices, and server-side runtimes. By treating local storage as the primary source of truth, it enables applications to remain functional without network connectivity, automatically reconciling changes with remote backends once a connection is restored. The database distinguishes itself through a modular architecture that supports cross-environment synchronization and high-performance d

    Combines high-speed memory storage with background disk replication to optimize query performance while ensuring long-term data durability.

    TypeScriptangularbrowser-databasecouchdb
    在 GitHub 上查看↗23,048
  • vectordotdev/vectorvectordotdev 的头像

    vectordotdev/vector

    22,071在 GitHub 上查看↗

    Vector is a high-performance observability data pipeline designed to collect, transform, and route logs, metrics, and traces across distributed infrastructure. It functions as a modular engine that decouples data ingestion from processing and transmission, utilizing a component-based architecture to connect diverse sources to multiple destinations. The project distinguishes itself through a focus on reliability and flow control. It implements backpressure-aware data movement to prevent data loss during traffic spikes and utilizes disk-backed event buffering to ensure durability during network

    Routes overflow events between memory and disk storage layers to balance speed and durability.

    Rusteventsforwarderhacktoberfest
    在 GitHub 上查看↗22,071
  • prestodb/prestoprestodb 的头像

    prestodb/presto

    16,711在 GitHub 上查看↗

    Presto is a distributed SQL query engine designed for high-performance analytical processing across heterogeneous data sources. It functions as a data federation platform and massively parallel processing engine, allowing users to execute interactive queries against diverse storage systems without requiring data migration. By mapping remote metadata and structures to a unified relational namespace, it enables seamless cross-platform analysis through a standard SQL interface. The engine distinguishes itself through a pluggable connector architecture and a shared-nothing distributed processing

    Offloads intermediate query results to disk when memory thresholds are exceeded to ensure stability during large-scale analytical processing.

    Javabig-datadatahadoop
    在 GitHub 上查看↗16,711
  • spotify/annoyspotify 的头像

    spotify/annoy

    14,157在 GitHub 上查看↗

    Annoy is a C++ library designed for approximate nearest neighbor search in high-dimensional vector spaces. It functions as a vector similarity search engine that constructs static, disk-based data structures to facilitate fast lookups. By mapping identifiers to vector data and persisting these structures to disk, the library enables efficient, memory-mapped access to large datasets. The project distinguishes itself through the use of random projection trees and distance-metric-based partitioning, which organize data into hierarchical binary trees to balance search precision against computatio

    Persists search structures to disk, allowing large datasets to be shared across processes without requiring full memory residency.

    C++approximate-nearest-neighbor-searchc-plus-plusgolang
    在 GitHub 上查看↗14,157
  • dask/daskdask 的头像

    dask/dask

    13,746在 GitHub 上查看↗

    Dask 是一个并行计算框架和分布式任务调度器,旨在将 Python 数据科学工作流从单机扩展到大型集群。它作为一个集群资源管理器,通过将任务及其依赖项表示为有向无环图来编排计算逻辑。这种架构允许系统在管理复杂执行要求的同时,自动将工作负载分配到可用硬件上。 该项目通过一个延迟评估引擎脱颖而出,该引擎将数据操作推迟到明确请求时才执行,从而实现全局图优化和高效的资源分配。它结合了内存感知数据溢出功能,以防止在处理超过可用内存的数据集时系统崩溃,并利用任务图融合将操作序列组合成单个执行步骤,从而最大限度地减少调度开销和节点间通信。 该平台为大规模数据分析提供了全面的功能面,包括对分布式机器学习、高性能计算集成和并行数据处理的支持。它提供了用于集群生命周期管理、性能分析和任务执行实时监控的广泛工具。用户可以在各种基础设施上部署这些环境,包括本地硬件、云提供商、容器化系统和高性能计算集群。

    Monitors memory usage during computation and offloads intermediate results to disk to prevent system crashes.

    Pythondasknumpypandas
    在 GitHub 上查看↗13,746
  • openrefine/openrefineOpenRefine 的头像

    OpenRefine/OpenRefine

    11,866在 GitHub 上查看↗

    OpenRefine is a data cleaning tool and wrangling platform used to transform raw, messy datasets into consistent and structured formats. It operates as a Java-based data processor that runs a local server and provides a web browser interface for managing and manipulating data. The platform includes a data reconciliation engine for matching local entries against external knowledge bases to standardize entities. It also functions as a web data augmentation tool, allowing users to fetch and integrate information from external web sources to enrich their datasets. The system provides a transforma

    Implements memory-to-disk spillover to handle datasets that exceed available RAM.

    Javadata-analysisdata-sciencedata-wrangling
    在 GitHub 上查看↗11,866
  • microsoft/fastermicrosoft 的头像

    microsoft/FASTER

    6,606在 GitHub 上查看↗

    FASTER is a high-throughput key-value store that combines an in-memory data store with a hybrid memory-disk storage engine, enabling datasets larger than available RAM. It uses a latch-free, cache-optimized index for concurrent point lookups and heavy updates, and records all mutations to a persistent append-only log on disk with checksum validation and group-commit checkpointing for crash recovery. The system supports multi-key transactional workloads through atomic multi-key locking, ensuring transactional consistency without coarse-grained contention. It exposes the key-value store to remo

    Keeps hot data in memory and seamlessly spills cold data to fast local or cloud storage.

    C#concurrenthash-tableindexing
    在 GitHub 上查看↗6,606
  • materializeinc/materializeMaterializeInc 的头像

    MaterializeInc/materialize

    6,314在 GitHub 上查看↗

    Materialize is a streaming SQL database that continuously ingests live data from sources such as Kafka, Redpanda, PostgreSQL, and MySQL, and incrementally maintains materialized views. It provides a PostgreSQL-compatible query engine that accepts standard SQL over the PostgreSQL wire protocol, enabling any existing SQL client or BI tool to query real-time data. The system also includes a Model Context Protocol (MCP) server that exposes live materialized view data to AI agents, providing fresh context without polling. Materialize distinguishes itself through its ability to offer configurable c

    Offloads large key-value state to disk automatically when using upsert or Debezium envelopes.

    Rust
    在 GitHub 上查看↗6,314
  • openatomfoundation/pikiwidbOpenAtomFoundation 的头像

    OpenAtomFoundation/pikiwidb

    6,113在 GitHub 上查看↗

    PikiwiDB is a distributed NoSQL database and disk-based key-value store that serves as a Redis-compatible protocol server. It is designed to handle datasets larger than available system memory by utilizing a persistence engine that stores the full dataset on disk. The system employs a tiered storage model, caching frequently accessed hot data in memory while maintaining the primary volume on disk. It ensures high availability through a replicated data store architecture, using asynchronous binary logs to synchronize data between primary and secondary nodes. The project supports distributed d

    Employs a tiered storage model that layers in-memory caching over a persistent disk-based storage engine.

    C++nosqlnosql-data-storagenosql-databases
    在 GitHub 上查看↗6,113
  • jerrylead/sparkinternalsJerryLead 的头像

    JerryLead/SparkInternals

    5,363在 GitHub 上查看↗

    SparkInternals 是一份技术参考和架构指南,详细介绍了 Apache Spark 分布式计算引擎的内部设计和实现。它作为大数据引擎分析的研究资料,重点关注系统如何管理集群执行以及驱动节点(Driver)、执行器(Executor)和工作节点(Worker)之间的交互。 该项目详细分解了逻辑计划如何转换为物理执行阶段。它专门分析了数据 Shuffle 操作、内存管理以及分布式作业调度协调的机制。 该文档涵盖了广泛的分布式计算功能,包括查询执行规划、数据依赖管理和内存缓存策略。它还研究了任务分配、并行执行以及用于故障恢复和数据持久化的过程。

    Offloads sorted key-value pairs to local disk when internal memory limits are exceeded during shuffles.

    在 GitHub 上查看↗5,363
  • trinea/android-commonTrinea 的头像

    Trinea/android-common

    5,022在 GitHub 上查看↗

    android-common is a collection of shared utility components and framework libraries for Android development. It provides specialized toolkits for reverse engineering, system utility management, data caching, and high-performance user interface components. The project includes a reverse engineering toolkit for inspecting application internals through package decompilation and manifest data extraction. It also features a system utility toolkit for managing file operations and executing shell commands within the Android operating system. The library covers several capability areas, including da

    Balances volatile memory for fast access with persistent disk storage for long-term asset retention.

    Java
    在 GitHub 上查看↗5,022
  • oceanbase/minioboceanbase 的头像

    oceanbase/miniob

    4,318在 GitHub 上查看↗

    MiniOB is an open-source educational relational database kernel designed for learning the internals of database systems. It implements a dual-engine storage architecture combining B+ Tree and LSM-Tree, supports SQL parsing and query execution, and provides transactional processing with multi-version concurrency control. The system communicates with clients using the MySQL wire protocol and includes a vector database extension for storing and querying high-dimensional vectors. The project distinguishes itself through its comprehensive coverage of core database concepts in a single, learnable c

    Manages disk page loading into a fixed-size memory pool with eviction for efficient access.

    C++classroomcplusplusdatabase
    在 GitHub 上查看↗4,318
  • facebookincubator/veloxfacebookincubator 的头像

    facebookincubator/velox

    4,155在 GitHub 上查看↗

    Velox 是一个高性能 C++ 查询执行引擎和列式数据处理库。它作为一个用于实现分析型查询引擎的可组合框架,提供了向量化表达式评估器和数据管理系统工具包。 该项目以使用向量化列式执行和基于 Arena 的内存分配来处理大规模数据集而著称。它具有专门的优化功能,如广播连接表缓存、动态过滤器下推和字典编码,以减少内存开销并加速分析读取。 该引擎涵盖了广泛的分析能力,包括实现哈希连接、合并连接和半连接,以及多阶段并行聚合和窗口函数计算。它提供了用于列式内存存储、Parquet 数据解码以及与云存储集成的原语。 通过用于自定义标量和聚合函数的函数注册系统提供可扩展性,并提供高级绑定以将 C++ 逻辑连接到 Python。

    Controls memory resource usage via arena allocation and spills intermediate data to disk when limits are exceeded.

    C++
    在 GitHub 上查看↗4,155
  • kuzudb/kuzukuzudb 的头像

    kuzudb/kuzu

    3,965在 GitHub 上查看↗

    Kùzu is an embedded property graph database engine designed for high-performance analytical queries and local data management. It operates as a library within the host application process, utilizing a columnar-based storage architecture and just-in-time query compilation to execute complex graph traversals and pattern matching efficiently. By mapping database files directly into system memory, it ensures data durability and high-speed access while maintaining ACID-compliant transactional integrity. The engine distinguishes itself by integrating vector similarity search and full-text search di

    Offloads intermediate query results to temporary disk storage when memory limits are reached to ensure processing stability.

    C++cypherdatabaseembeddable
    在 GitHub 上查看↗3,965
  1. Home
  2. Data & Databases
  3. Persistent Storage Providers
  4. Memory-Disk Layering

探索子标签

  • Buffer Pool Page Evictions1 个子标签Loads disk pages into a fixed-size memory frame pool and evicts old pages when space runs out. **Distinct from Memory-Disk Layering:** Distinct from Memory-Disk Layering: focuses on the buffer pool eviction policy for database pages, not general memory-disk layering.
  • Hot-Cold Data SpillingStorage layers that automatically move cold data from memory to disk or cloud storage while keeping hot data in memory. **Distinct from Memory-Disk Layering:** Distinct from Memory-Disk Layering: focuses on the automatic spilling of cold data to external storage, not just the general layering of memory over disk.
  • Memory-Spilling EnginesExecution engines that offload intermediate query results to disk when memory thresholds are exceeded. **Distinct from Memory-Disk Layering:** Distinct from general memory-disk layering: focuses on the execution engine's stability mechanism during large-scale processing.
  • Multi-Layered Buffer TopologiesArchitectures that route data between memory and disk storage layers to balance performance and durability. **Distinct from Memory-Disk Layering:** Distinct from Memory-Disk Layering: focuses on the topology of chaining buffers for overflow management rather than general storage layering.