Kubernetes incidents, explained
See what broke.Understand why.Know what to do next.
kwatch turns Kubernetes failures into clear alerts with the likely cause, useful evidence, and a practical next step.
OOMKilled
production / orders-api
Pod: orders-api-7ffc9d4f9-x9p4t · Node: worker-3
The container exceeded its 512Mi memory limit.
Increase limits.memory or reduce memory usage.
Recent logs and Kubernetes events add context to the alert.
From signal to action
Understand the incident without piecing it together yourself
Kubernetes shows symptoms. kwatch connects the story so your team can decide what needs attention.
Detect the problem
Watch for crashes, stuck workloads, unhealthy nodes, and other cluster signals.
Connect the clues
Bring together status, recent logs, Kubernetes events, and affected resources.
Send a useful alert
Give responders a likely cause and a next step in the channel they already use.
Coverage
Start with common failures. Add checks as you grow.
Safe defaults cover everyday incidents. Heartbeat, Metrics Server usage, TLS checks, and active probes are available when you need them.
Pods and scheduling
Crashes, OOM kills, restarts, readiness, and pending Pods.
Workloads
Rollouts, Jobs, CronJobs, autoscaling, and availability.
Infrastructure and storage
Node pressure, persistent storage, and platform health.
Networking and security
Services, Ingress, webhooks, TLS, RBAC, and policy findings.
Get started
From command to useful alerts.
The interactive kwatch.sh manager guides installation, channel setup, and verification. You need Bash, curl, kubectl, and cluster install permissions.
Run the manager on a machine with access to your cluster.
/bin/bash -c "$(curl -fsSL https://kwatch.dev/kwatch.sh)"
Run it again to change settings, upgrade, check status, or uninstall.
Choose your cluster
The manager shows the current kubectl context before installing.
Connect a channel
It stores credentials in a Kubernetes Secret.
Verify the install
It checks the deployment before you start monitoring.
Two replicas by default: one active leader and one standby. A single-replica option is available without kwatch self-failover.
How failover worksNotifications
Send alerts where your team works.
Connect a familiar destination first. Choose from 56 integrations when your team needs more routes.
See all channels