awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

12 个仓库

Awesome GitHub RepositoriesData Archiving

Patterns for long-term storage and preservation of data streams.

Distinguishing note: Focuses on messaging stream archiving patterns.

Explore 12 awesome GitHub repositories matching software engineering & architecture · Data Archiving. Refine with filters or upvote what's useful.

Awesome Data Archiving GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • nsqio/nsqnsqio 的头像

    nsqio/nsq

    25,738在 GitHub 上查看↗

    NSQ is a distributed, brokerless messaging platform designed for high-throughput, fault-tolerant communication. By utilizing a decentralized topology, it eliminates single points of failure and allows for horizontal scaling across clusters. The system organizes message streams into topics and channels, effectively decoupling producers from consumers to support both streaming and job-oriented workloads. The platform distinguishes itself through a lookup-service-based discovery mechanism that enables clients to dynamically locate producers at runtime without requiring centralized coordination.

    Archives message streams to disk using dedicated consumer channels.

    Godistributed-systemsgomessage-queue
    在 GitHub 上查看↗25,738
  • letta-ai/lettaletta-ai 的头像

    letta-ai/letta

    21,168在 GitHub 上查看↗

    Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across long-term interactions. It provides a comprehensive suite of primitives for defining agents with configurable personas, modular memory blocks, and tool-use capabilities, enabling them to retain user preferences and conversation history over extended sessions. The platform distinguishes itself through its advanced memory management and orchestration capabilities. It allows agents to autonomously update their own memory, perform retrieval-augmented generation, and coordinate com

    Adds text content to persistent archives with vector embeddings for semantic retrieval.

    Pythonaiai-agentsllm
    在 GitHub 上查看↗21,168
  • microsoftdocs/azure-docsMicrosoftDocs 的头像

    MicrosoftDocs/azure-docs

    10,894在 GitHub 上查看↗

    Azure Docs is the official technical documentation repository for Microsoft Azure, the cloud computing platform. It provides comprehensive guidance on the full spectrum of Azure services, covering everything from core infrastructure components like virtual machines, Kubernetes clusters, and serverless computing to platform services for AI, machine learning, data analytics, and storage. The documentation details how to provision, manage, and govern cloud resources at scale, including policy enforcement, identity management, and cost optimization. The documentation distinguishes Azure through i

    Documents Azure's archival storage service for rarely used data at the lowest cost.

    Markdownskilling
    在 GitHub 上查看↗10,894
  • blinkospace/blinkoblinkospace 的头像

    blinkospace/blinko

    10,601在 GitHub 上查看↗

    Blinko is a personal knowledge management system and an LLM-powered knowledge base that enables users to capture and organize thoughts through a bi-directional knowledge graph. It functions as a RAG-enabled note-taking application and a self-hosted Markdown editor, allowing for the creation of permanent documentation and fleeting notes. The project distinguishes itself by integrating retrieval-augmented generation to provide conversational querying and AI-powered analysis of private document libraries. It supports both cloud-based and local AI model integration, enabling users to perform sema

    Automatically moves old records into an archive at specified frequencies to maintain workspace cleanliness.

    TypeScriptmarkdownmemosnextjs
    在 GitHub 上查看↗10,601
  • tyrrrz/discordchatexporterTyrrrz 的头像

    Tyrrrz/DiscordChatExporter

    10,392在 GitHub 上查看↗

    Discord Chat Exporter is a tool for extracting messages and media from Discord channels and direct messages into offline files. It functions as a backup utility and archival tool, using authentication tokens to retrieve chat history and metadata for long-term storage or history recovery. The system converts API data into readable documents and supports multi-format export options, including HTML, TXT, CSV, and JSON. It includes capabilities for automated chat backups by creating recurring tasks through the host operating system's task scheduler. The tool provides data management features suc

    Saves chat history as HTML, JSON, CSV, or TXT files for permanent long-term storage.

    C#archivalarchiverchat
    在 GitHub 上查看↗10,392
  • 1n3/sn1per1N3 的头像

    1N3/Sn1per

    10,049在 GitHub 上查看↗

    Sn1per is a vulnerability management platform and penetration testing orchestrator designed to automate reconnaissance, vulnerability scanning, and exploit verification. It functions as a dockerized security toolkit that coordinates multiple tools into a unified automated pipeline to identify security flaws across network and web assets. The platform features an attack surface manager for discovering internet-facing assets through OSINT, DNS enumeration, and certificate transparency. It distinguishes itself with an AI-powered security analyzer that uses large language models to summarize scan

    Bundles workspace reports, loot, and vulnerability files into a single compressed archive for backup.

    Shellattack-surfaceattack-surface-managementattacksurface
    在 GitHub 上查看↗10,049
  • supabase/realtimesupabase 的头像

    supabase/realtime

    7,488在 GitHub 上查看↗

    Realtime is a real-time data distribution and synchronization engine that enables applications to stream database changes and coordinate state between clients. It functions as a synchronization layer that monitors database write-ahead logs to provide change data capture and pushes updates to authorized clients via WebSockets. The project features a real-time presence server for tracking the online status of active users and a broadcast service for sending ephemeral messages without database persistence. It organizes communication through channel-based message routing and uses a structured JSO

    Stores processed messages in an archive to maintain audit trails and reference history.

    Elixircdcchange-data-capturecrdt
    在 GitHub 上查看↗7,488
  • infobyte/faradayinfobyte 的头像

    infobyte/faraday

    6,523在 GitHub 上查看↗

    Faraday is a vulnerability management platform and security tool aggregator designed to centralize security findings from multiple scanners into a single dashboard. It utilizes a relational security database to catalog hosts, services, and security flaws, enabling users to track remediation and analyze organizational risk. The platform distinguishes itself through a plugin-based system that normalizes diverse security tool outputs into a unified data model. It supports deep integration with a wide array of scanners and CLI tools, intercepting shell command output or parsing report files to ag

    Allows hiding completed or inactive workspaces to maintain focus on active security engagements.

    Python
    在 GitHub 上查看↗6,523
  • j3ssie/osmedeusj3ssie 的头像

    j3ssie/Osmedeus

    6,425在 GitHub 上查看↗

    Osmedeus is a security workflow orchestration engine that coordinates AI agents, shell commands, and scanning tools through declarative YAML pipelines. It functions as a distributed security scanner, a declarative workflow automator, and an AI agent framework for security, enabling automated multi-step security analysis with conditional branching, parallel execution, and distributed workers. The engine distinguishes itself through a hybrid runner model that executes workflow steps on the local host, inside Docker containers, or over SSH to remote machines, selected per step or module. It supp

    Packages workspace files, database records, and metadata into a ZIP archive for backup or transfer.

    Go
    在 GitHub 上查看↗6,425
  • owid/covid-19-dataowid 的头像

    owid/covid-19-data

    5,663在 GitHub 上查看↗

    Data on COVID-19 (coronavirus) cases, deaths, hospitalizations, tests • All countries • Updated daily by Our World in Data

    A structured archive of global COVID-19 indicators with consistent column metadata and versioned releases.

    Pythoncoronaviruscovidcovid-19
    在 GitHub 上查看↗5,663
  • yacy/yacy_search_serveryacy 的头像

    yacy/yacy_search_server

    3,966在 GitHub 上查看↗

    yacysearchserver 是一个去中心化索引系统和点对点搜索引擎。它作为分布式网络爬虫和内联网搜索设备运行,允许在不依赖中央机构的情况下发现和索引网络内容。 该项目支持创建用于索引内部网站和文件系统的私有搜索门户。它利用点对点网络共享搜索索引,并将查询路由分发到多个服务器实例,通过去中心化的索引管理来扩展搜索结果。 该系统涵盖递归网络爬取、网页内容索引和高级查询优化。它包括用于数据备份、恢复以及索引数据导出和导入的工具。 该应用程序支持通过 Docker 容器部署,并允许通过环境变量进行系统配置。

    Preserves system state and search indices by serializing data volumes into compressed backup files.

    Javadecentralizedintranet-searchintranet-search-engine
    在 GitHub 上查看↗3,966
  • swe-agent/mini-swe-agentSWE-agent 的头像

    SWE-agent/mini-swe-agent

    2,947在 GitHub 上查看↗

    mini-swe-agent is an autonomous software engineering system designed to develop features and fix bugs by combining large language models with a bash interface. It operates as an agentic framework that executes coding tasks and documentation updates through a continuous cycle of model reasoning and tool execution. The project differentiates itself with a strong focus on safety and evaluation, utilizing container-based sandbox execution via Docker or Singularity to isolate command execution. It includes a batch-parallel evaluation harness to measure code-fixing accuracy against standardized sof

    Archives the state of the containerized workspace into local tarballs to preserve task outcomes.

    Pythonagentagentic-aiagentic-ai-cli
    在 GitHub 上查看↗2,947
  1. Home
  2. Software Engineering & Architecture
  3. Data Archiving

探索子标签

  • Epidemiological Data ArchivesStructured archives of global health indicators with consistent column metadata and versioned releases. **Distinct from Data Archiving:** Distinct from Data Archiving: focuses on public health epidemiological data with metadata rather than general messaging stream archiving.
  • Infrequent Data ArchivalStores rarely used data at the lowest cost with retrieval times measured in hours. **Distinct from Data Archiving:** Distinct from Data Archiving: focuses on low-cost archival storage for infrequently accessed data, not general archiving patterns.
  • Workspace Archiving1 个子标签Bundling and compressing engagement-specific data, reports, and loot for backup. **Distinct from Data Archiving:** Specifically archives penetration testing workspace data rather than generic system data streams.