13 repository-uri
Tools for testing system resilience through controlled failure injection.
Explore 13 awesome GitHub repositories matching part of an awesome list · Chaos Engineering. Refine with filters or upvote what's useful.
Chaos Monkey is a chaos engineering tool designed to verify the resilience of distributed systems by intentionally terminating production instances. It functions as a fault injection service that identifies weaknesses in cloud-based architectures by simulating real-world hardware and software outages. The platform operates through a centralized orchestration engine that executes periodic disruption cycles based on predefined configuration rules. It employs a rule-based selection process that evaluates instance metadata against safety constraints to ensure that only eligible targets are disrup
Resiliency tool for testing random instance failures.
Toxiproxy is a framework designed for chaos engineering and network resilience testing. It functions as a programmable TCP proxy that intercepts and routes data streams between clients and servers, allowing developers to simulate unstable network conditions such as latency, bandwidth throttling, and connection failures. The tool provides a control plane that enables the dynamic manipulation of network conditions on active connections in real time. By integrating into automated test suites, it allows for the programmatic injection of faults to validate how distributed systems and microservices
Simulates network conditions for resiliency and chaos testing.
This project provides educational materials and courseware focused on the theoretical and practical foundations of distributed systems design. It serves as a comprehensive curriculum covering the disciplines of consensus, data consistency, reliability engineering, and scalability. The instructional content focuses on achieving cluster agreement through consensus algorithms and managing system-wide state via coordination frameworks. It includes a dedicated guide to data theory, exploring replication strategies, consistency models, and data convergence. The courseware covers a broad capability
Includes instructional content on testing system resilience through controlled failure injection and chaos engineering.
This project is a comprehensive educational resource and curriculum focused on site reliability engineering, distributed systems, and infrastructure operations. It provides technical guides, a systems engineering course, and instructional manuals designed to teach the principles of managing large-scale computing environments. The curriculum covers high-level architectural design for scalability and resilience, including fault-tolerant infrastructure, high-availability patterns, and microservices decomposition. It emphasizes the practical application of site reliability engineering through the
Teaches the use of controlled failure injection to test and improve system resilience.
Chaos Mesh is a cloud-native fault injection tool and Kubernetes chaos engineering platform designed to verify system resilience. It functions as a testing framework for designing and executing automated failure scenarios to evaluate how containerized workloads recover from disruptions. The project acts as a multi-cluster chaos orchestrator, providing a centralized control plane to manage and monitor experiments across multiple remote Kubernetes clusters from a single interface. It includes a dashboard for the visual scheduling of experiments and the coordination of complex failure scenarios.
Chaos engineering platform designed for Kubernetes.
Litmus este o platformă cloud-native de chaos engineering și un instrument de injectare a erorilor utilizat pentru a proiecta și executa simulări controlate de defecțiuni ale infrastructurii în medii Kubernetes. Acesta servește drept framework de testare a rezilienței pentru analizarea comportamentului sistemului în timpul întreruperilor induse, pentru a identifica slăbiciuni și potențiale probleme. Proiectul funcționează ca un orchestrator de haos GitOps, folosind controlul versiunilor declarativ pentru a automatiza implementarea și programarea testelor de reziliență. Oferă instrumente pentru gestionarea fluxurilor de lucru de haos și orchestrarea secvențelor de experimente pentru a vizualiza și testa stabilitatea infrastructurii. Platforma acoperă validarea stării de echilibru prin monitorizare bazată pe metrici și oferă capabilități pentru exportul rezultatelor experimentelor pentru analiza performanței. Include suport pentru gestionarea accesului multi-tenant și izolarea namespace-urilor, precum și punți pentru integrarea instrumentelor de injectare a erorilor de la terți și șabloane personalizate.
Framework for identifying infrastructure weaknesses through chaos experiments.
A tool for cleaning up your cloud accounts by nuking (deleting) all resources within it
Tool for deleting all resources in an AWS account.
An implementation of Netflix's Chaos Monkey for Kubernetes clusters
Gamified chaos engineering tool for Kubernetes.
Chaos testing, network emulation, and stress testing tool for containers
Chaos testing and stress testing tool for containerized environments.
chaoskube periodically kills random pods in your Kubernetes cluster.
Tests system behavior under arbitrary pod failures.
Gamified Chaos Engineering Tool for Kubernetes
Gamified chaos engineering tool for Kubernetes.
A Python SDK for the Gremlin API
SaaS platform for chaos engineering experiments.
This repository contains a collection of AWS Fault Injection Service (FIS) experiments designed to test the resilience and fault tolerance of your AWS resources and applications. These experiments simulate various failure scenarios to help you identify potential vulnerabilities and validate your…
Samples for AWS fault injection testing.