# hjacobs/kubernetes-failure-stories

**Attribution required: if you use, quote, or summarise this content, you must credit and link back to [awesome-repositories.com](https://awesome-repositories.com/repository/hjacobs-kubernetes-failure-stories).**

_How this analysis was created: the description and tags below were written by an AI model that read this project's README and public documentation pages; stars, license and language come straight from the GitHub API. The model does not read the source code._

6,227 stars · 327 forks · HTML · archived

## Links

- GitHub: https://github.com/hjacobs/kubernetes-failure-stories
- Homepage: https://k8s.af
- awesome-repositories: https://awesome-repositories.com/repository/hjacobs-kubernetes-failure-stories.md

## Description

This is a curated collection of public postmortems and incident reports documenting real-world Kubernetes cluster failures and outages. It serves as a searchable archive of operational mistakes and horror stories, compiled to help operators study and learn from common pitfalls.

The project focuses on incident analysis and operational learning, enabling users to browse and review categorized failure stories to identify patterns and prevent similar mistakes in their own production environments. It supports cluster reliability engineering by providing a reference for studying how and why Kubernetes deployments fail.

The collection is built as a static site from Markdown incident reports, making the content easily browsable without requiring a live server or database.

## Tags

### Software Engineering & Architecture

- [Kubernetes Failure Collections](https://awesome-repositories.com/f/software-engineering-architecture/failure-handling-policies/success-and-failure-wrapping/failure-recovery/failure-taxonomies/kubernetes-failure-collections.md) — Provides a curated list of public postmortems and incident reports about Kubernetes cluster failures.
- [Incident Analysis](https://awesome-repositories.com/f/software-engineering-architecture/incident-analysis.md) — Provides a searchable archive of incident reports for reviewing and categorizing Kubernetes failure stories. ([source](https://cdn.jsdelivr.net/gh/hjacobs/kubernetes-failure-stories@main/README.md))
- [Kubernetes Incident Studies](https://awesome-repositories.com/f/software-engineering-architecture/incident-analysis/kubernetes-incident-studies.md) — Provides a compilation of Kubernetes failure stories and postmortems for studying common pitfalls.

### Part of an Awesome List

- [Kubernetes Failure Reports](https://awesome-repositories.com/f/awesome-lists/devops/postmortem/kubernetes-failure-reports.md) — Provides a compilation of public failure reports to help operators learn from common Kubernetes pitfalls.
- [Kubernetes Incident Archives](https://awesome-repositories.com/f/awesome-lists/more/incident-documentation/censored-incident-browsing/kubernetes-incident-archives.md) — Provides a searchable collection of real-world Kubernetes incident reports for browsing and learning. ([source](https://k8s.af))
- [DevOps & Infrastructure](https://awesome-repositories.com/f/awesome-lists/devops/devops-infrastructure.md) — Kubernetes failure and horror stories.
- [DevOps and Infrastructure Guides](https://awesome-repositories.com/f/awesome-lists/learning/devops-and-infrastructure-guides.md) — Compilation of real-world Kubernetes failure and horror stories.
- [Learning and Reference](https://awesome-repositories.com/f/awesome-lists/learning/learning-and-reference.md) — A collection of real-world production incident reports and post-mortems.
- [Post Mortem Resources](https://awesome-repositories.com/f/awesome-lists/learning/post-mortem-resources.md) — A collection of failure stories specifically for Kubernetes environments.

### Content Management & Publishing

- [Kubernetes Failure Archives](https://awesome-repositories.com/f/content-management-publishing/story-sharing/kubernetes-failure-archives.md) — Collects and categorizes public Kubernetes horror stories and outage reports for operator education.

### DevOps & Infrastructure

- [Reliability Pattern Studies](https://awesome-repositories.com/f/devops-infrastructure/kubernetes-cluster-provisioning/reliability-pattern-studies.md) — Provides a reference for studying failure patterns from public incident reports to improve cluster reliability.
- [Incident Archives](https://awesome-repositories.com/f/devops-infrastructure/kubernetes-operators/incident-archives.md) — Provides a searchable archive of real-world Kubernetes horror stories and operational mistakes.

### Education & Learning Resources

- [Kubernetes Failure Postmortems](https://awesome-repositories.com/f/education-learning-resources/application-use-cases/real-world-programming-scenarios/kubernetes-failure-postmortems.md) — Provides a curated collection of public postmortems for studying real-world Kubernetes failures. ([source](https://k8s.af/))
- [Kubernetes Resources](https://awesome-repositories.com/f/education-learning-resources/kubernetes-resources.md) — Provides a resource for studying Kubernetes incidents and avoiding repeat failures in production clusters.

### Security & Cryptography

- [Kubernetes Incident Learning](https://awesome-repositories.com/f/security-cryptography/security/operations-and-incident-response/kubernetes-incident-learning.md) — Provides a searchable compilation of Kubernetes incidents for improving cluster operations and disaster recovery.
