18 repository-uri
Deployment strategies that ensure cluster-wide availability through redundant control planes and agent nodes.
Distinct from High-Availability Deployments: Distinct from High-Availability Deployments: focuses on the entire cluster architecture rather than specific Helm chart replicas.
Explore 18 awesome GitHub repositories matching devops & infrastructure · High Availability Cluster Deployments. Refine with filters or upvote what's useful.
This project is a local Kubernetes cluster manager and tool that runs control plane and worker nodes as containers on a host machine. It provides an environment for local development and automated testing by emulating a full Kubernetes cluster within a container runtime. The tool enables the creation of multi-node topologies and high-availability control planes through configuration files. It supports image sideloading to transfer container images directly from the host to nodes, bypassing remote registries, and allows for offline deployments using pre-built node images. Capabilities include
Defines and deploys complex Kubernetes architectures with multiple control planes and worker nodes for high availability testing.
vcluster is a Kubernetes virtual cluster platform that creates fully isolated Kubernetes environments with dedicated control planes, API servers, and RBAC on shared physical infrastructure. It virtualizes Kubernetes control planes by running them as pods inside a host cluster, as standalone binaries on bare metal or virtual machines, or within Docker containers, providing each tenant their own isolated Kubernetes environment without the overhead of managing separate physical clusters. The platform enables multi-tenant Kubernetes isolation through multiple tenancy models, from shared node pool
Runs virtual cluster control planes with multiple replicas and failover for production resilience.
This project provides a comprehensive guide and set of scripts for deploying and configuring a production-ready Kubernetes cluster from scratch. It centers on establishing a functional environment by installing core management components, storage, and networking across multiple nodes. The implementation emphasizes high availability for the control plane, utilizing layer-4 load balancing and leader election for the API server, scheduler, and controller manager. It further ensures reliability through the deployment of a distributed key-value store for persistent runtime data. The project cover
Ensures API server availability by deploying multiple instances behind a layer-4 load balancer.
SpringCloud-Learning is an educational project that demonstrates how to build microservices using Spring Cloud, covering the core patterns of service discovery, centralized configuration management, and API gateway routing. The project provides hands-on examples for registering and discovering microservice instances with Nacos, Eureka, or Consul, and for routing external API requests through Spring Cloud Gateway with support for filters and load balancing. The tutorials explore service resilience through circuit breakers and rate limiting with Sentinel and Hystrix, including custom fallback l
Demonstrates deploying multiple Nacos configuration server instances for high availability.
This project is a collection of learning resources and instructional guides for implementing asynchronous messaging patterns using RabbitMQ. It provides a series of tutorials and runnable code examples focused on the Advanced Message Queuing Protocol to help users decouple services via a message broker. The resources cover practical implementation patterns including request-reply, pub-sub, and stream processing. These guides demonstrate how to use official client libraries to balance worker loads, route messages across multiple consumers in a distributed system, and deploy high availability b
Provides documentation and examples for deploying high availability broker clusters on Kubernetes.
k3d is a containerized Kubernetes orchestrator and wrapper that manages the lifecycle of k3s nodes and servers within container runtimes. It serves as a tool for deploying and destroying multi-node Kubernetes environments on a single machine for local development and testing. The project distinguishes itself by wrapping k3s to provide integrated networking, resource limit control, and local registry orchestration. It enables multi-node cluster simulation by wrapping nodes as isolated containers and using host-entry injection and port proxying to route host TCP and UDP traffic into the cluster
Emulates multi-node cluster topologies and high availability configurations on a single host using containers.
k0s is a Kubernetes distribution that packages all control plane and worker components into a single binary, enabling cluster deployment with no host dependencies beyond the Linux kernel. It supports a container-native runtime where controllers and workers run inside Docker containers using a single OCI image, and offers declarative YAML configuration for defining cluster topology, host roles, and SSH connection details. The distribution provides pre-built binaries for x86-64, ARM64, ARMv7, and RISC-V architectures, and uses cryptographically signed tokens for secure node enrollment. The proj
Expands a cluster from one node to a multi-node setup with isolated control planes for fault tolerance.
gpt-load is a transparent proxy gateway that routes API requests to multiple AI providers—including OpenAI, Google Gemini, and Anthropic Claude—through a single endpoint while preserving each provider's native format and authentication. It acts as a centralized routing layer, allowing applications to switch between AI services by changing only the base URL without modifying any client code or business logic. The proxy distinguishes itself through intelligent traffic management across pools of API keys, offering automatic key rotation, weighted or round-robin load balancing, and failover that
Deploying a master-slave architecture with horizontal scaling, configuration sync, and graceful shutdown for enterprise use.
pfSense is an open-source operating system that turns a standard computer into a dedicated firewall and router with a web-based management interface. It runs on the FreeBSD kernel with the pf packet filter for stateful firewall and NAT processing, and manages all settings through a PHP-driven web interface that writes to XML configuration files. The platform provides a comprehensive set of network security capabilities accessible through its browser-based control panel. Users can configure packet filtering rules to control traffic flow between network segments, manage network address translat
Pairs two firewalls in a failover cluster using CARP for seamless high availability.
m3 is a distributed time series database designed for high-resolution metrics and high-cardinality data management. It functions as a scalable storage system and a multi-cluster query engine, providing a distributed metrics aggregator capable of downsampling and summarizing data before it is committed to storage. The project distinguishes itself through a coordinated cluster model using etcd for node membership and shard placement. It supports multiple ingestion protocols, including the Prometheus remote write protocol, InfluxDB line protocol, and Graphite Carbon plaintext protocol, and provi
Uses leader and follower replication to ensure data flushing continues if a primary node fails.
Pomerium este un reverse proxy de tip zero-trust, conștient de identitate, conceput pentru a securiza aplicațiile interne și infrastructura fără a fi nevoie de un VPN tradițional. Prin verificarea fiecărei cereri printr-un furnizor de identitate și un motor de politici declarativ, acesta asigură că accesul este acordat pe baza identității utilizatorului, contextului dispozitivului și atributelor cererii, mai degrabă decât pe locația rețelei. Proiectul se distinge prin capacitatea sa de a gestiona protocoale diverse și cerințe complexe de acces. Oferă acces securizat, verificat prin identitate, la aplicații web, servicii TCP și UDP și conexiuni SSH, înlocuind cheile statice cu certificate efemere, cu durată scurtă de viață. În plus, funcționează ca un gateway pentru Model Context Protocol, permițând descoperirea securizată, auditarea și execuția de instrumente impusă de politici pentru agenții AI și serviciile automatizate. Pomerium se integrează cu Kubernetes ca un ingress controller, traducând resursele native de ingress în configurații de rutare conștiente de identitate. Arhitectura sa suportă sincronizarea centralizată a planului de control, permițând administratorilor să distribuie politici de securitate, certificate TLS și configurații de rutare în clustere proxy distribuite. Sistemul oferă, de asemenea, gestionarea granulară a traficului, inclusiv load balancing, manipularea header-elor și filtrarea bazată pe path, pentru a asigura o conectivitate securizată și performantă către serviciile backend. Proiectul este conceput pentru un deployment flexibil, suportând atât arhitecturi unificate, cât și decuplate pentru a acomoda scalarea independentă a componentelor de date și de control plane. Oferă o observabilitate cuprinzătoare prin audit logging-ul deciziilor de autorizare și exportul metricilor de performanță, facilitând monitorizarea conformității și securității în medii diverse.
Deploys and scales local proxy instances connected to a centralized control plane for unified policy distribution.
caddy-docker-proxy este un reverse proxy HTTP dinamic și un controller de ingress pentru rețeaua Docker care generează automat configurații de rutare prin citirea etichetelor din containerele Docker. Acesta servește ca instrument de service discovery care detectează adresele IP ale containerelor în timp real pentru a ruta traficul web primit către țintele backend corecte. Proiectul funcționează ca un orchestrator de proxy distribuit, capabil să împingă configurațiile generate de la un controller central către mai multe instanțe de server remote pentru a scala gestionarea cererilor. Automatizează emiterea și reînnoirea certificatelor de securitate TLS pentru domeniile proxate și coordonează certificatele partajate între replicile de server. Sistemul suportă generarea de configurații bazată pe etichete, utilizând template-uri și Docker secrets pentru a injecta metadate și setări globale în configurația serverului. Menține sincronizarea prin ascultarea evenimentelor din ciclul de viață al containerelor și utilizează un mecanism de încărcare in-process pentru a actualiza comportamentul serverului. Capabilitățile de gestionare a traficului includ definirea regulilor de rutare a cererilor, potrivirea căilor și rescrierea URI-urilor pentru a direcționa traficul către resursele containerizate.
Deploys multiple server instances coordinated by a central controller to scale request handling.
Mgmt este un sistem distribuit de gestionare a configurațiilor care menține starea dorită a clusterelor folosind automatizarea bazată pe evenimente și feedback-ul closed-loop. Funcționează ca un motor de automatizare a infrastructurii care declanșează corecții ale stării sistemului în timp real, pe baza monitorizării resurselor și a specificațiilor predefinite. Sistemul include un selector de noduri de cluster distribuit pentru alegerea subseturilor de gazde pe baza unor strategii și constrângeri specifice pentru a distribui sarcinile de lucru. Dispune, de asemenea, de un manager de infrastructură cloud pentru controlul ciclului de viață al instanțelor de mașini virtuale, inclusiv deployment-ul imaginilor, selecția regiunii și scripturile de startup, alături de un orchestrator de firmware hardware pentru instalarea și verificarea binarilor pe baseboard management controllers (BMC). Capabilitățile suplimentare acoperă execuția sarcinilor în paralel, modificarea fișierelor de configurare bazată pe lens și regăsirea metadatelor sistemului. Setul de instrumente oferă, de asemenea, utilitare pentru transformarea tipurilor de date, decodarea formatelor și inspecția mediului pentru a determina stările curente ale sistemului.
Ships a distributed node selector to elect host subsets based on specific strategies and constraints.
Emitter is a distributed pub-sub platform and message brokering system that decouples senders and receivers through topic-based routing. It functions as an MQTT message broker and WebSocket communication server, enabling real-time data exchange between hardware devices and web-based clients. The system acts as a secure channel orchestrator and message persistence engine. It ensures delivery through a storage system that buffers historical messages and queues data for offline subscribers. Access to these data streams is managed via time-limited permission keys that enforce granular read and wr
Supports deploying synchronized node networks to increase throughput and ensure continuous service availability.
This project is a Terraform Kubernetes provisioner and K3s cluster deployer designed to automate the installation and configuration of lightweight container orchestration on Hetzner Cloud infrastructure. It functions as a Hetzner Cloud infrastructure module, using declarative configuration to manage the full lifecycle of virtual machines, private networks, and load balancers. The orchestrator focuses on high availability by deploying redundant control planes and worker nodes across multiple physical data centers to ensure service continuity. It incorporates a cloud network security manager to
Supports redundant control planes and agent nodes deployed across multiple zones for fault tolerance.
Light Task Scheduler is a distributed job scheduling and workflow orchestration platform designed for managing background processing across scalable computing environments. It functions as a cluster management system that coordinates stateless nodes to execute recurring, cron-based, or one-time tasks with centralized control and high availability. The platform distinguishes itself through a leader-based coordination model that automatically elects a primary controller to manage task distribution and system state. It supports complex workflow dependencies, ensuring that prerequisite tasks comp
Provides a platform for managing distributed computing nodes with automatic master election and high availability.
Ansible Interactive Tutorial is a learning platform and command-line training utility that runs interactive configuration management exercises inside pre-configured Docker containers. It provides an isolated multi-container sandbox environment featuring a control node and target hosts configured with pre-shared SSH keys for practicing infrastructure automation step by step. The platform guides users through structured, sequential learning paths from fundamental to advanced configuration management concepts. It supports volume-mounted local workspaces that bind host directories directly into r
Deploys a multi-node topology emulating isolated runtime instances with pre-configured networking and security trust relationships.
XWiki Platform is a collaborative content management system and enterprise wiki designed for creating, organizing, and sharing structured documentation. It functions as a Java-based application framework that enables teams to build data-driven business applications directly within a web environment. By combining a flexible knowledge base with modular development tools, the platform supports both standard document management and the creation of custom, interactive software solutions. The platform distinguishes itself through a highly extensible architecture that allows for deep customization w
Distributes traffic across multiple server instances connected to a shared database to balance load and improve performance.